Model lifecycle

GPT-5.6 Terra

Released 2026-07-10 · Back to GPT

Exact-version signals observed 2026-08-13–2026-08-24 · n=4

Collecting a stable baseline A verdict appears after the first 14-day window reaches n=30 across at least 7 active days.

Lifecycle updated daily · latest complete UTC day 2026-08-27

Fixed 14-day baseline
Latest 28 days
n=4
Change vs baseline
90-day slope
pts / 30 days

14-day rolling experience

Each daily point summarizes the previous 14 complete UTC days. The dashed line is the fixed release baseline; the verdict still uses independent non-overlapping 14-day segments.

Waiting for a sample-sufficient baseline

A complete window needs n=30 across at least 7 active days · current signals n=4

What explains the change

Experience change and discussion-mix change are separated. These are observational contributions, not proof of cause.

No category has enough samples in both windows for a contribution conclusion.

Category Baseline → current Category change Weight share Experience contribution Discussion-mix contribution
Coding Baseline → currentnot enough datan=0 → 0 Category change Weight share Experience contribution Discussion-mix contribution
Image / vision Baseline → currentnot enough datan=0 → 0 Category change Weight share Experience contribution Discussion-mix contribution
Local deploy Baseline → currentnot enough datan=0 → 0 Category change Weight share Experience contribution Discussion-mix contribution
Reasoning Baseline → currentnot enough datan=0 → 1 Category change Weight share Experience contribution Discussion-mix contribution
Roleplay / creative Baseline → currentnot enough datan=0 → 0 Category change Weight share Experience contribution Discussion-mix contribution
Safety & refusals Baseline → currentnot enough datan=0 → 0 Category change Weight share Experience contribution Discussion-mix contribution
Speed & latency Baseline → currentnot enough datan=0 → 1 Category change Weight share Experience contribution Discussion-mix contribution
General text Baseline → currentnot enough datan=0 → 2 Category change Weight share Experience contribution Discussion-mix contribution
Video generation Baseline → currentnot enough datan=0 → 0 Category change Weight share Experience contribution Discussion-mix contribution

Lifecycle scores use the same experience-signal weights as the main index, without sample shrinkage after n=30. They measure public user perception, not model capability or backend causes.