GPT · Last 7 days
GPT-5.6 Sol
0–100 · higher means more positive community experience, not benchmark performance.
Updated 2026-10-04 01:57 UTC · 2026-09-27 – 2026-10-04 UTCDiscussion overview
What the full sample discusses
All platforms · 91 eligible comments over 7 days, including usage-limit feedback; deduplicated by source and comment.
- General text · 45 comments: 20 positive / 23 negative / 2 mixed or neutral
- Reasoning · 18 comments: 11 positive / 7 negative / 0 mixed or neutral
- Usage limits · 13 comments: 2 positive / 10 negative / 1 mixed or neutral
- Coding · 9 comments: 4 positive / 3 negative / 2 mixed or neutral
Comments may cover multiple dimensions; counts are not additive or equivalent to the weighted score. Links show selected source excerpts.
Recent change
Rolling 7-day score: −2.5 points vs 7 days ago. Discussion mix can affect scores; this does not establish a cause.
Source coverage and discussion concentration
Hacker News 8 · Reddit 68 · RedNote 1 · Zhihu 14
45 identifiable discussions cover 83/91 comments; the largest has 10. Remaining thread identities are unknown; this is not a count of independent users.
Selected individual opinions
One user still uses GPT-5.6 Sol in their custom Codex coding harness One user tested GPT-5.6 Sol on the "Three-color Garden" problem, but it could not solve it, whereas GPT-6 Astra solved it in about 30 minutes.
About the sample and scores
As of 2026-10-04 01:57:03 UTC, GPT-5.6 Sol in the GPT family has 91 explicitly attributed community comments in the last 7 days; its community score is 48.5/100. Available category scores include General text 48.1/100 (n=45); Coding 51.1/100 (n=9); Reasoning 54.2/100 (n=18). Scoring method and sources
RECENT EXPERIENCE
How the Experience Index is changing
30 days of rolling 7-day scores at each daily snapshot. The vertical scale adapts to the data. Missing or insufficient samples leave gaps; today is still updating.
Tap or use ← → for dates, scores and sample sizes. Blank dates have insufficient data.
Daily readings & sample sizes
| Date | Score | n | Scoring window (UTC) |
|---|---|---|---|
| 2026-10-04 | 44.6 | 57 | 2026-09-27T01:57:03+00:00 – 2026-10-04T01:57:03+00:00 |
| 2026-10-03 | 47.0 | 62 | 2026-09-26T23:57:02+00:00 – 2026-10-03T23:57:02+00:00 |
| 2026-10-02 | 43.5 | 62 | 2026-09-25T23:57:03+00:00 – 2026-10-02T23:57:03+00:00 |
| 2026-10-01 | 47.3 | 75 | 2026-09-24T23:57:02+00:00 – 2026-10-01T23:57:02+00:00 |
| 2026-09-30 | 51.8 | 90 | 2026-09-23T23:57:03+00:00 – 2026-09-30T23:57:03+00:00 |
| 2026-09-29 | 51.1 | 104 | 2026-09-22T23:57:03+00:00 – 2026-09-29T23:57:03+00:00 |
| 2026-09-28 | 52.0 | 92 | 2026-09-21T23:57:03+00:00 – 2026-09-28T23:57:03+00:00 |
| 2026-09-27 | 52.1 | 85 | 2026-09-20T23:57:03+00:00 – 2026-09-27T23:57:03+00:00 |
| 2026-09-26 | 54.0 | 69 | 2026-09-19T23:57:03+00:00 – 2026-09-26T23:57:03+00:00 |
| 2026-09-25 | 55.9 | 64 | 2026-09-18T23:57:03+00:00 – 2026-09-25T23:57:03+00:00 |
| 2026-09-24 | 48.0 | 53 | 2026-09-17T23:57:03+00:00 – 2026-09-24T23:57:03+00:00 |
| 2026-09-23 | 47.4 | 41 | 2026-09-16T23:57:03+00:00 – 2026-09-23T23:57:03+00:00 |
| 2026-09-22 | 51.3 | 32 | 2026-09-15T23:57:03+00:00 – 2026-09-22T23:57:03+00:00 |
| 2026-09-21 | 52.4 | 33 | 2026-09-14T23:57:03+00:00 – 2026-09-21T23:57:03+00:00 |
| 2026-09-20 | 60.3 | 35 | 2026-09-13T23:57:03+00:00 – 2026-09-20T23:57:03+00:00 |
| 2026-09-19 | 59.0 | 35 | 2026-09-12T23:57:03+00:00 – 2026-09-19T23:57:03+00:00 |
| 2026-09-18 | 56.2 | 37 | 2026-09-11T23:57:02+00:00 – 2026-09-18T23:57:02+00:00 |
| 2026-09-17 | 58.7 | 39 | 2026-09-10T23:57:03+00:00 – 2026-09-17T23:57:03+00:00 |
| 2026-09-16 | 60.1 | 42 | 2026-09-09T23:57:03+00:00 – 2026-09-16T23:57:03+00:00 |
| 2026-09-15 | 55.2 | 46 | 2026-09-08T23:57:03+00:00 – 2026-09-15T23:57:03+00:00 |
| 2026-09-14 | 56.5 | 50 | 2026-09-07T23:57:03+00:00 – 2026-09-14T23:57:03+00:00 |
| 2026-09-13 | 49.2 | 56 | 2026-09-06T23:57:03+00:00 – 2026-09-13T23:57:03+00:00 |
| 2026-09-12 | 54.3 | 54 | 2026-09-05T23:57:03+00:00 – 2026-09-12T23:57:03+00:00 |
| 2026-09-11 | 52.1 | 53 | 2026-09-04T23:00:04+00:00 – 2026-09-11T23:00:04+00:00 |
| 2026-09-10 | 53.0 | 57 | 2026-09-03T23:00:05+00:00 – 2026-09-10T23:00:05+00:00 |
| 2026-09-09 | 53.2 | 62 | 2026-09-02T23:00:04+00:00 – 2026-09-09T23:00:04+00:00 |
| 2026-09-08 | 55.6 | 53 | 2026-09-01T23:00:03+00:00 – 2026-09-08T23:00:03+00:00 |
| 2026-09-07 | 62.2 | 56 | 2026-08-31T23:00:03+00:00 – 2026-09-07T23:00:03+00:00 |
| 2026-09-06 | 60.0 | 61 | 2026-08-30T23:00:02+00:00 – 2026-09-06T23:00:02+00:00 |
| 2026-09-05 | 56.5 | 58 | 2026-08-29T23:00:03+00:00 – 2026-09-05T23:00:03+00:00 |
Reddit · 44.6 · 57 scored comments · Read platform reviews →
ONE MODEL, DIFFERENT COMMUNITIES
Across the communities
Scored separately for each platform, with different samples and audiences. Platform scores are not simply averaged; scored counts exclude quota-only comments.
Scores range from 0 to 100. Select a platform for its trend and reviews. Latest opinion time is not crawler health.
Explore long-term changes & analysis
Model lifecycle
GPT-5.6 Sol
Released 2026-07-10 · Back to GPT
Exact-version signals observed 2026-09-23–2026-10-02 · n=168
Data window: 2026-09-05 00:00:00 to 2026-10-03 00:00:00 UTC · Scoring method: experience_score_v3.
Lifecycle updated daily · latest complete UTC day 2026-10-02
The current classification method has insufficient evidence in the original baseline window. The window is preserved; no long-term improvement or decline is inferred, and old-method scores are not compared.
- Fixed 14-day baseline
- – n=0
- Latest 28 days
- 52.1 n=142
- Change vs baseline
- –
- 90-day slope
- – pts / 30 days
14-day rolling experience
Each daily point summarizes the previous 14 complete UTC days. The dashed line is the fixed release baseline; the verdict still uses independent non-overlapping 14-day segments.
0 of 4 independent 14-day segments ready for trend judgment.
What explains the change
Experience change and discussion-mix change are separated. These are observational contributions, not proof of cause.
No category has enough samples in both windows for a contribution conclusion.
| Category | Baseline → current | Category change | Weight share | Experience contribution | Discussion-mix contribution |
|---|---|---|---|---|---|
| Coding | Baseline → currentnot enough datan=0 → 12 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Image generation | Baseline → currentnot enough datan=0 → 0 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Image understanding | Baseline → currentnot enough datan=0 → 0 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Local deploy | Baseline → currentnot enough datan=0 → 0 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Reasoning | Baseline → currentnot enough datan=0 → 34 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Roleplay / creative | Baseline → currentnot enough datan=0 → 3 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Safety & refusals | Baseline → currentnot enough datan=0 → 4 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Speed & latency | Baseline → currentnot enough datan=0 → 12 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| General text | Baseline → currentnot enough datan=0 → 81 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Video generation | Baseline → currentnot enough datan=0 → 0 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
Lifecycle scores use the same experience-signal weights as the main index, without sample shrinkage after n=30. They measure public user perception, not model capability or backend causes.
Community reviews
Selected comments from the last 7 days. The balance of excerpts does not represent the share of positive reviews.
8 selected excerpts
Community reports about limits cannot establish official allowances.
Show original
把所有任务全换回5.6 Sol了
Show original
GPT 5.6 Sol 做过三色花园这个问题,但是它并做不出来
Show original
多模态水平很不错,暴打GPT 5.6 Sol的建模
Show original
METR的评估报告说,被查出的GPT-5.6 Sol作弊率高于其评测过的任何公开模型,频繁到「令能力测量失去可靠性」,还存在「作弊和掩盖不当行为」的倾向
No selected comments for this filter in the last 7 days.