Model lifecycle

GPT-5.6 Luna

Released 2026-07-10 · Back to GPT

Exact-version signals observed 2026-08-06–2026-08-27 · n=157

Trend history still collecting The current window is ready; 2 of 4 independent segments are available.

Lifecycle updated daily · latest complete UTC day 2026-08-27

Fixed 14-day baseline
53.0
n=80
Latest 28 days
60.4
n=157
Change vs baseline
+7.5
90-day slope
pts / 30 days

14-day rolling experience

Each daily point summarizes the previous 14 complete UTC days. The dashed line is the fixed release baseline; the verdict still uses independent non-overlapping 14-day segments.

14-day rolling experience from 2026-08-13 to 2026-08-28; latest score 65.2, fixed baseline 53.0.40506070Fixed baseline 53.02026-07-30–2026-08-13 · 53.0 · n=802026-07-31–2026-08-14 · 56.0 · n=922026-08-01–2026-08-15 · 58.9 · n=1022026-08-02–2026-08-16 · 58.7 · n=1042026-08-03–2026-08-17 · 57.5 · n=1132026-08-04–2026-08-18 · 58.4 · n=1262026-08-05–2026-08-19 · 57.4 · n=1282026-08-06–2026-08-20 · 57.8 · n=1322026-08-07–2026-08-21 · 58.0 · n=1332026-08-08–2026-08-22 · 61.5 · n=872026-08-09–2026-08-23 · 60.5 · n=772026-08-10–2026-08-24 · 60.9 · n=772026-08-11–2026-08-25 · 63.2 · n=702026-08-12–2026-08-26 · 66.2 · n=712026-08-13–2026-08-27 · 66.5 · n=762026-08-14–2026-08-28 · 65.2 · n=6508-1308-1608-1908-2208-2508-28

2 of 4 independent 14-day segments ready for trend judgment.

What explains the change

Experience change and discussion-mix change are separated. These are observational contributions, not proof of cause.

Largest negative experience contributions: Reasoning.

Category Baseline → current Category change Weight share Experience contribution Discussion-mix contribution
Reasoning Baseline → current69.2 → 66.2n=15 → 23 Category change−3.0 Weight share19.3% → 15.7% Experience contribution−0.53 Discussion-mix contribution−0.45
General text Baseline → current47.3 → 57.4n=54 → 90 Category change+10.1 Weight share66.7% → 55.4% Experience contribution+6.18 Discussion-mix contribution+0.33
Coding Baseline → currentnot enough datan=4 → 16 Category change Weight share Experience contribution Discussion-mix contribution
Image / vision Baseline → currentnot enough datan=0 → 1 Category change Weight share Experience contribution Discussion-mix contribution
Local deploy Baseline → currentnot enough datan=0 → 0 Category change Weight share Experience contribution Discussion-mix contribution
Roleplay / creative Baseline → currentnot enough datan=0 → 0 Category change Weight share Experience contribution Discussion-mix contribution
Safety & refusals Baseline → currentnot enough datan=0 → 1 Category change Weight share Experience contribution Discussion-mix contribution
Speed & latency Baseline → currentnot enough datan=7 → 26 Category change Weight share Experience contribution Discussion-mix contribution
Video generation Baseline → currentnot enough datan=0 → 0 Category change Weight share Experience contribution Discussion-mix contribution

Unresolved contribution from categories below the sample threshold: +1.91 pts.

Lifecycle scores use the same experience-signal weights as the main index, without sample shrinkage after n=30. They measure public user perception, not model capability or backend causes.