Model lifecycle
GPT-5.6 Sol
Released 2026-07-10 · Back to GPT
Exact-version signals observed 2026-08-07–2026-08-27 · n=401
Lifecycle updated daily · latest complete UTC day 2026-08-27
- Fixed 14-day baseline
- 56.1 n=120
- Latest 28 days
- 49.5 n=401
- Change vs baseline
- −6.6
- 90-day slope
- – pts / 30 days
14-day rolling experience
Each daily point summarizes the previous 14 complete UTC days. The dashed line is the fixed release baseline; the verdict still uses independent non-overlapping 14-day segments.
2 of 4 independent 14-day segments ready for trend judgment.
What explains the change
Experience change and discussion-mix change are separated. These are observational contributions, not proof of cause.
Largest negative experience contributions: Coding, General text, Speed & latency.
| Category | Baseline → current | Category change | Weight share | Experience contribution | Discussion-mix contribution |
|---|---|---|---|---|---|
| Coding | Baseline → current52.9 → 43.7n=26 → 76 | Category change−9.1 | Weight share22.5% → 19.7% | Experience contribution−1.92 | Discussion-mix contribution+0.12 |
| General text | Baseline → current45.6 → 39.9n=33 → 150 | Category change−5.8 | Weight share26.4% → 36.4% | Experience contribution−1.81 | Discussion-mix contribution−0.97 |
| Speed & latency | Baseline → current64.8 → 57.7n=25 → 69 | Category change−7.1 | Weight share19.1% → 16.6% | Experience contribution−1.27 | Discussion-mix contribution−0.21 |
| Reasoning | Baseline → current61.1 → 65.9n=32 → 82 | Category change+4.8 | Weight share28.9% → 21.5% | Experience contribution+1.20 | Discussion-mix contribution−0.82 |
| Image / vision | Baseline → currentnot enough datan=0 → 16 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Local deploy | Baseline → currentnot enough datan=0 → 0 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Roleplay / creative | Baseline → currentnot enough datan=0 → 0 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Safety & refusals | Baseline → currentnot enough datan=4 → 8 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Video generation | Baseline → currentnot enough datan=0 → 0 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
Unresolved contribution from categories below the sample threshold: −0.95 pts.
Lifecycle scores use the same experience-signal weights as the main index, without sample shrinkage after n=30. They measure public user perception, not model capability or backend causes.