Claude · 近 7 天
Claude Sonnet 5
0–100 分,越高代表社区使用体验越正面,不代表能力测试成绩。
更新于 2026-10-02 15:22 UTC · 2026-09-25 – 2026-10-02 UTC全部平台体感分32.5/10030 条评论 · 近 7 天
大家怎么看
有用户发现,Opus 5.5 在小型任务上会走捷径,而 Sonnet 和 Fable 本可以正确完成这些任务。 有用户连续使用 codex/GPT 5.6 一个月,整体体验明显优于 Claude Sonnet 5。
样本与评分说明
截至 2026-10-02 15:22:03 UTC,Claude 家族的 Claude Sonnet 5 在近 7 天有 30 条明确归属该版本的社区反馈,社区口碑分为 32.5/100。部分有效分类评分:通用文本 38.4/100 (n=19);速度与延迟 34.9/100 (n=5)。 评分方法与数据来源
近期走势
口碑在如何变化
历史不足较 7 天前 · 分
近 30 天,每点代表截至该日采样时刻的近 7 天滚动体感分。纵轴随数据调整;缺失或样本不足处断开,今天仍在更新。
点按或用 ← → 查看日期、分数与样本量;空白日期暂无足够数据。
查看每日读数与样本量
| 日期 | 体感分 | n | 评分窗口(UTC) |
|---|---|---|---|
| 2026-10-02 | 30.7 | 8 | 2026-09-25T15:22:03+00:00 – 2026-10-02T15:22:03+00:00 |
| 2026-10-01 | 30.7 | 8 | 2026-09-24T23:57:02+00:00 – 2026-10-01T23:57:02+00:00 |
| 2026-09-30 | 30.7 | 8 | 2026-09-23T23:57:03+00:00 – 2026-09-30T23:57:03+00:00 |
| 2026-09-29 | 31.7 | 9 | 2026-09-22T23:57:03+00:00 – 2026-09-29T23:57:03+00:00 |
| 2026-09-28 | 31.8 | 9 | 2026-09-21T23:57:03+00:00 – 2026-09-28T23:57:03+00:00 |
| 2026-09-27 | — | 3 | 2026-09-20T23:57:03+00:00 – 2026-09-27T23:57:03+00:00 |
| 2026-09-26 | — | 2 | 2026-09-19T23:57:03+00:00 – 2026-09-26T23:57:03+00:00 |
| 2026-09-25 | — | 2 | 2026-09-18T23:57:03+00:00 – 2026-09-25T23:57:03+00:00 |
| 2026-09-24 | — | 2 | 2026-09-17T23:57:03+00:00 – 2026-09-24T23:57:03+00:00 |
| 2026-09-23 | — | 1 | 2026-09-16T23:57:03+00:00 – 2026-09-23T23:57:03+00:00 |
| 2026-09-22 | — | 2 | 2026-09-15T23:57:03+00:00 – 2026-09-22T23:57:03+00:00 |
| 2026-09-21 | — | 2 | 2026-09-14T23:57:03+00:00 – 2026-09-21T23:57:03+00:00 |
| 2026-09-20 | — | 2 | 2026-09-13T23:57:03+00:00 – 2026-09-20T23:57:03+00:00 |
| 2026-09-19 | — | 2 | 2026-09-12T23:57:03+00:00 – 2026-09-19T23:57:03+00:00 |
| 2026-09-18 | — | 2 | 2026-09-11T23:57:02+00:00 – 2026-09-18T23:57:02+00:00 |
| 2026-09-17 | — | 3 | 2026-09-10T23:57:03+00:00 – 2026-09-17T23:57:03+00:00 |
| 2026-09-16 | — | 3 | 2026-09-09T23:57:03+00:00 – 2026-09-16T23:57:03+00:00 |
| 2026-09-15 | — | 2 | 2026-09-08T23:57:03+00:00 – 2026-09-15T23:57:03+00:00 |
| 2026-09-14 | — | 1 | 2026-09-07T23:57:03+00:00 – 2026-09-14T23:57:03+00:00 |
| 2026-09-13 | — | 1 | 2026-09-06T23:57:03+00:00 – 2026-09-13T23:57:03+00:00 |
| 2026-09-12 | — | 3 | 2026-09-05T23:57:03+00:00 – 2026-09-12T23:57:03+00:00 |
| 2026-09-11 | — | 3 | 2026-09-04T23:00:04+00:00 – 2026-09-11T23:00:04+00:00 |
| 2026-09-10 | — | 4 | 2026-09-03T23:00:05+00:00 – 2026-09-10T23:00:05+00:00 |
| 2026-09-09 | 58.5 | 6 | 2026-09-02T23:00:04+00:00 – 2026-09-09T23:00:04+00:00 |
| 2026-09-08 | 50.7 | 8 | 2026-09-01T23:00:03+00:00 – 2026-09-08T23:00:03+00:00 |
| 2026-09-07 | 50.7 | 8 | 2026-08-31T23:00:03+00:00 – 2026-09-07T23:00:03+00:00 |
| 2026-09-06 | 46.9 | 7 | 2026-08-30T23:00:02+00:00 – 2026-09-06T23:00:02+00:00 |
| 2026-09-05 | 42.6 | 6 | 2026-08-29T23:00:03+00:00 – 2026-09-05T23:00:03+00:00 |
| 2026-09-04 | 38.4 | 5 | 2026-08-28T23:00:03+00:00 – 2026-09-04T23:00:03+00:00 |
| 2026-09-03 | — | 4 | 2026-08-27T23:00:03+00:00 – 2026-09-03T23:00:03+00:00 |
Hacker News · 30.7 · 8 条评分评论 · 查看该平台评价 →
不同社区,同一模型
各平台怎么看
各平台独立计算,样本量和讨论人群不同。平台分数不取简单平均;评分评论不包含仅讨论额度的评论。
Reddit最新评分评论 2026-10-01 01:04 UTC38.419 条较7天前 −7.7Hacker News最新评分评论 2026-09-28 23:29 UTC30.78 条较7天前 历史不足知乎最新评分评论 2026-09-30 00:35 UTC—2 条样本不足小红书暂无评分评论时间—0 条样本不足
分数范围 0–100。点击平台查看趋势及对应评价。最新评论时间不代表爬虫运行状态。
社区评论
精选近 7 天的不同意见,摘录条数不代表真实好评比例。
4 条精选摘录
额度评价来自社区体验,不能据此推算官方实际配额。
展开原文
it was so much better than opus or sonnet 5
展开原文
sonnet 5 was basically unusable
展开原文
both Sonnet or Fable would have nailed every time
查看评论上下文
但在几个任务之后,我意识到它遗漏了复杂任务的大量细节,而Fable绝不会遗漏这些。它反而创建了一个能通过自己测试的版本,并断言工作已完成(公平地说,所有LLM都会这样做)。但与5.1相比,它就像是在故意走捷径来让工作"完成"。看起来它被调校成会修剪任务中"非必要"的部分,以便更快地到达终点。 而且我不认为这仅限于那些有解释空间的大型任务。我发现在小得多的附带任务上它也会这样做,而Sonnet或Fable每次都能完美完成这些任务。
机器翻译展开原文
But after a few tasks I realised it was omitting huge details of complex tasks that Fable just would not have missed. It instead creates a version that passes its own tests and asserts that the job is done (which to be fair, all LLMs do). But in comparison to 5.1, its like it takes intentional shortcuts to get the job 'done'. It looks like it's been tuned to trim "non-essential" aspects of a task to get to the end faster. And I don't think this is isolated to large tasks that have room for interpretation. I've caught doing the same with much smaller side tasks that both Sonnet or Fable would have nailed every time.
展开原文
Sonnet 5 was so horrible
展开原文
Sonnet 5 to perform worse across all of our evals, especially against time
展开原文
I would never get an answer back even for very simple prompts. It would just churn on nothing and return max token usage reached
展开原文
Sonnet 5 is genuinely a terrible model for me
展开原文
how obnoxious and lazy Sonnet 5 was
展开原文
It does worse than Sonnet 5. Mainly because it is more reluctant to keep going to get an answer
查看评论上下文
我构建了一种对抗性的深奥编程语言来对LLM模型进行基准测试,刚在Sonnet 5.5上运行了它。它的表现比Sonnet 5还差。
机器翻译展开原文
I built an adversarial esoteric programming language to benchmark LLM models and just ran it on Sonnet 5.5 It does worse than Sonnet 5.
展开原文
I even replaced it with sonnet 5 recently because I wasn't very impressed with the results
查看评论上下文
在工作中我们有 copilot,即使有 750 美元的额度,我也很少使用 opus,我最近甚至用 sonnet 5 替换了它,因为我对结果不太满意,等不及下周我们启用 opus 5.5 了。
机器翻译展开原文
At work we have copilot and even with 750$ credits I use opus very sparingly and I even replaced it with sonnet 5 recently because I wasn't very impressed with the results, can't wait we enable opus 5.5 next week.
展开原文
It told me it had made everything up. So, I asked it to search on actual websites. It kept making things up.
查看评论上下文
我让它查找一些工业自动化产品的在线价格。
机器翻译展开原文
I asked it to find online prices for some industrial automation items.
近 7 天暂时没有符合此筛选条件的精选评论。