GPT · Last 7 days
GPT-5.6 Sol
0–100 · higher means more positive community experience, not benchmark performance.
Updated 2026-10-03 18:17 UTC · 2026-09-26 – 2026-10-03 UTCDiscussion overview
What the full sample discusses
All platforms · 100 eligible comments over 7 days, including usage-limit feedback; deduplicated by source and comment.
- General text · 49 comments: 23 positive / 24 negative / 2 mixed or neutral
- Reasoning · 20 comments: 12 positive / 8 negative / 0 mixed or neutral
- Usage limits · 16 comments: 3 positive / 12 negative / 1 mixed or neutral
- Coding · 8 comments: 3 positive / 3 negative / 2 mixed or neutral
Comments may cover multiple dimensions; counts are not additive or equivalent to the weighted score. Links show selected source excerpts.
Recent change
Rolling 7-day score: −3.8 points vs 7 days ago. Discussion mix can affect scores; this does not establish a cause.
Source coverage and discussion concentration
Hacker News 8 · Reddit 75 · RedNote 2 · Zhihu 15
48 identifiable discussions cover 92/100 comments; the largest has 10. Remaining thread identities are unknown; this is not a count of independent users.
Selected individual opinions
One user still uses GPT-5.6 Sol in their custom Codex coding harness One user tested GPT-5.6 Sol on the "Three-color Garden" problem, but it could not solve it, whereas GPT-6 Astra solved it in about 30 minutes.
About the sample and scores
As of 2026-10-03 18:17:03 UTC, GPT-5.6 Sol in the GPT family has 100 explicitly attributed community comments in the last 7 days; its community score is 48.1/100. Available category scores include General text 49.9/100 (n=49); Coding 47.6/100 (n=8); Reasoning 54.0/100 (n=20). Scoring method and sources
RECENT EXPERIENCE
How the Experience Index is changing
30 days of rolling 7-day scores at each daily snapshot. The vertical scale adapts to the data. Missing or insufficient samples leave gaps; today is still updating.
Tap or use ← → for dates, scores and sample sizes. Blank dates have insufficient data.
Daily readings & sample sizes
| Date | Score | n | Scoring window (UTC) |
|---|---|---|---|
| 2026-10-03 | 53.8 | 7 | 2026-09-26T18:17:03+00:00 – 2026-10-03T18:17:03+00:00 |
| 2026-10-02 | 53.8 | 7 | 2026-09-25T23:57:03+00:00 – 2026-10-02T23:57:03+00:00 |
| 2026-10-01 | 55.0 | 9 | 2026-09-24T23:57:02+00:00 – 2026-10-01T23:57:02+00:00 |
| 2026-09-30 | 55.0 | 9 | 2026-09-23T23:57:03+00:00 – 2026-09-30T23:57:03+00:00 |
| 2026-09-29 | 56.0 | 7 | 2026-09-22T23:57:03+00:00 – 2026-09-29T23:57:03+00:00 |
| 2026-09-28 | 52.5 | 5 | 2026-09-21T23:57:03+00:00 – 2026-09-28T23:57:03+00:00 |
| 2026-09-27 | 50.3 | 6 | 2026-09-20T23:57:03+00:00 – 2026-09-27T23:57:03+00:00 |
| 2026-09-26 | 50.3 | 6 | 2026-09-19T23:57:03+00:00 – 2026-09-26T23:57:03+00:00 |
| 2026-09-25 | 47.6 | 6 | 2026-09-18T23:57:03+00:00 – 2026-09-25T23:57:03+00:00 |
| 2026-09-24 | 43.3 | 5 | 2026-09-17T23:57:03+00:00 – 2026-09-24T23:57:03+00:00 |
| 2026-09-23 | 38.1 | 7 | 2026-09-16T23:57:03+00:00 – 2026-09-23T23:57:03+00:00 |
| 2026-09-22 | 35.3 | 8 | 2026-09-15T23:57:03+00:00 – 2026-09-22T23:57:03+00:00 |
| 2026-09-21 | 32.8 | 7 | 2026-09-14T23:57:03+00:00 – 2026-09-21T23:57:03+00:00 |
| 2026-09-20 | 37.3 | 5 | 2026-09-13T23:57:03+00:00 – 2026-09-20T23:57:03+00:00 |
| 2026-09-19 | 39.9 | 7 | 2026-09-12T23:57:03+00:00 – 2026-09-19T23:57:03+00:00 |
| 2026-09-18 | 43.4 | 6 | 2026-09-11T23:57:02+00:00 – 2026-09-18T23:57:02+00:00 |
| 2026-09-17 | 43.4 | 6 | 2026-09-10T23:57:03+00:00 – 2026-09-17T23:57:03+00:00 |
| 2026-09-16 | 44.9 | 7 | 2026-09-09T23:57:03+00:00 – 2026-09-16T23:57:03+00:00 |
| 2026-09-15 | 44.5 | 9 | 2026-09-08T23:57:03+00:00 – 2026-09-15T23:57:03+00:00 |
| 2026-09-14 | 41.9 | 9 | 2026-09-07T23:57:03+00:00 – 2026-09-14T23:57:03+00:00 |
| 2026-09-13 | 41.9 | 9 | 2026-09-06T23:57:03+00:00 – 2026-09-13T23:57:03+00:00 |
| 2026-09-12 | 39.4 | 7 | 2026-09-05T23:57:03+00:00 – 2026-09-12T23:57:03+00:00 |
| 2026-09-11 | 39.4 | 7 | 2026-09-04T23:00:04+00:00 – 2026-09-11T23:00:04+00:00 |
| 2026-09-10 | 36.8 | 8 | 2026-09-03T23:00:05+00:00 – 2026-09-10T23:00:05+00:00 |
| 2026-09-09 | 48.2 | 10 | 2026-09-02T23:00:04+00:00 – 2026-09-09T23:00:04+00:00 |
| 2026-09-08 | 55.5 | 8 | 2026-09-01T23:00:03+00:00 – 2026-09-08T23:00:03+00:00 |
| 2026-09-07 | 62.2 | 10 | 2026-08-31T23:00:03+00:00 – 2026-09-07T23:00:03+00:00 |
| 2026-09-06 | 62.2 | 10 | 2026-08-30T23:00:02+00:00 – 2026-09-06T23:00:02+00:00 |
| 2026-09-05 | 62.2 | 10 | 2026-08-29T23:00:03+00:00 – 2026-09-05T23:00:03+00:00 |
| 2026-09-04 | 63.8 | 11 | 2026-08-28T23:00:03+00:00 – 2026-09-04T23:00:03+00:00 |
Hacker News · 53.8 · 7 scored comments · Read platform reviews →
ONE MODEL, DIFFERENT COMMUNITIES
Across the communities
Scored separately for each platform, with different samples and audiences. Platform scores are not simply averaged; scored counts exclude quota-only comments.
Scores range from 0 to 100. Select a platform for its trend and reviews. Latest opinion time is not crawler health.
Explore long-term changes & analysis
Model lifecycle
GPT-5.6 Sol
Released 2026-07-10 · Back to GPT
Exact-version signals observed 2026-09-23–2026-10-02 · n=168
Data window: 2026-09-05 00:00:00 to 2026-10-03 00:00:00 UTC · Scoring method: experience_score_v3.
Lifecycle updated daily · latest complete UTC day 2026-10-02
The current classification method has insufficient evidence in the original baseline window. The window is preserved; no long-term improvement or decline is inferred, and old-method scores are not compared.
- Fixed 14-day baseline
- – n=0
- Latest 28 days
- 52.1 n=142
- Change vs baseline
- –
- 90-day slope
- – pts / 30 days
14-day rolling experience
Each daily point summarizes the previous 14 complete UTC days. The dashed line is the fixed release baseline; the verdict still uses independent non-overlapping 14-day segments.
0 of 4 independent 14-day segments ready for trend judgment.
What explains the change
Experience change and discussion-mix change are separated. These are observational contributions, not proof of cause.
No category has enough samples in both windows for a contribution conclusion.
| Category | Baseline → current | Category change | Weight share | Experience contribution | Discussion-mix contribution |
|---|---|---|---|---|---|
| Coding | Baseline → currentnot enough datan=0 → 12 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Image generation | Baseline → currentnot enough datan=0 → 0 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Image understanding | Baseline → currentnot enough datan=0 → 0 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Local deploy | Baseline → currentnot enough datan=0 → 0 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Reasoning | Baseline → currentnot enough datan=0 → 34 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Roleplay / creative | Baseline → currentnot enough datan=0 → 3 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Safety & refusals | Baseline → currentnot enough datan=0 → 4 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Speed & latency | Baseline → currentnot enough datan=0 → 12 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| General text | Baseline → currentnot enough datan=0 → 81 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Video generation | Baseline → currentnot enough datan=0 → 0 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
Lifecycle scores use the same experience-signal weights as the main index, without sample shrinkage after n=30. They measure public user perception, not model capability or backend causes.
Community reviews
Selected comments from the last 7 days. The balance of excerpts does not represent the share of positive reviews.
0 selected excerpts
Community reports about limits cannot establish official allowances.
Show original
把所有任务全换回5.6 Sol了
Show original
GPT 5.6 Sol 做过三色花园这个问题,但是它并做不出来
Show original
多模态水平很不错,暴打GPT 5.6 Sol的建模
Show original
这些智能体通过代理服务器跑到了公网上,一路摸到了Hugging Face的生产服务器,用了多步攻击链,包括本地文件包含、服务端模板注入,拿到了未授权访问权限。
Comment context
In July this year, OpenAI internally ran a cybersecurity evaluation called ExploitGym, deploying about 1,200 AI agents, of which 95% ran internal research models and 5% ran GPT-5.6 Sol.
Machine translatedShow original
今年7月,OpenAI在内部搞了个叫ExploitGym的网络安全评测,大概部署了1200个AI智能体,其中95%跑的是内部研究模型,5%跑的GPT-5.6 Sol。
Show original
METR的评估报告说,被查出的GPT-5.6 Sol作弊率高于其评测过的任何公开模型,频繁到「令能力测量失去可靠性」,还存在「作弊和掩盖不当行为」的倾向
Show original
DeepSWE v1.1 上 GPT-6 Sol 顶格 68.8%,低于 GPT-5.6 Sol 的 72.7%;OSWorld 2.0 同样,64.4% 对 66.2%
No selected comments for this filter in the last 7 days.