Model lifecycle
GPT-5.6 Luna
Released 2026-07-10 · Back to GPT
Exact-version signals observed 2026-08-06–2026-08-27 · n=157
Lifecycle updated daily · latest complete UTC day 2026-08-27
- Fixed 14-day baseline
- 53.0 n=80
- Latest 28 days
- 60.4 n=157
- Change vs baseline
- +7.5
- 90-day slope
- – pts / 30 days
14-day rolling experience
Each daily point summarizes the previous 14 complete UTC days. The dashed line is the fixed release baseline; the verdict still uses independent non-overlapping 14-day segments.
2 of 4 independent 14-day segments ready for trend judgment.
What explains the change
Experience change and discussion-mix change are separated. These are observational contributions, not proof of cause.
Largest negative experience contributions: Reasoning.
| Category | Baseline → current | Category change | Weight share | Experience contribution | Discussion-mix contribution |
|---|---|---|---|---|---|
| Reasoning | Baseline → current69.2 → 66.2n=15 → 23 | Category change−3.0 | Weight share19.3% → 15.7% | Experience contribution−0.53 | Discussion-mix contribution−0.45 |
| General text | Baseline → current47.3 → 57.4n=54 → 90 | Category change+10.1 | Weight share66.7% → 55.4% | Experience contribution+6.18 | Discussion-mix contribution+0.33 |
| Coding | Baseline → currentnot enough datan=4 → 16 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Image / vision | Baseline → currentnot enough datan=0 → 1 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Local deploy | Baseline → currentnot enough datan=0 → 0 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Roleplay / creative | Baseline → currentnot enough datan=0 → 0 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Safety & refusals | Baseline → currentnot enough datan=0 → 1 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Speed & latency | Baseline → currentnot enough datan=7 → 26 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Video generation | Baseline → currentnot enough datan=0 → 0 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
Unresolved contribution from categories below the sample threshold: +1.91 pts.
Lifecycle scores use the same experience-signal weights as the main index, without sample shrinkage after n=30. They measure public user perception, not model capability or backend causes.