Claude · 近 7 天
Claude Sonnet 5
0–100 分,越高代表社区使用体验越正面,不代表能力测试成绩。
更新于 2026-10-03 13:42 UTC · 2026-09-26 – 2026-10-03 UTC讨论概览
主要讨论什么
全部平台 · 近7天 28 条合格评论,包含额度反馈;按来源与评论去重。
- 通用文本 · 18 条:正面 6 / 负面 11 / 褒贬或中性 1
- 速度与延迟 · 6 条:正面 1 / 负面 4 / 褒贬或中性 1
- 推理 · 3 条:正面 1 / 负面 2 / 褒贬或中性 0
- 编程 · 2 条:正面 1 / 负面 1 / 褒贬或中性 0
同一评论可涉及多个维度,各项不可相加;方向计数不等于加权评分。链接展示该维度的精选原文。
近期变化
近7天滚动体感分较7天前 −12.2 分。讨论构成也会影响分数,这不是原因诊断。
样本分布与讨论集中度
Hacker News 8 · Reddit 18 · 知乎 2
可识别 18 个讨论,覆盖 20/28 条评论;最大单帖 2 条。其余讨论归属未知,不推算独立用户数。
个体观点选摘
有用户因6.1 Sol速度慢、沟通差,倾向选用Sonnet 5和Gemini Flash 3.8作为替代。 有用户连续使用 codex/GPT 5.6 一个月,整体体验明显优于 Claude Sonnet 5。
样本与评分说明
截至 2026-10-03 13:42:04 UTC,Claude 家族的 Claude Sonnet 5 在近 7 天有 28 条明确归属该版本的社区反馈,社区口碑分为 36.9/100。部分有效分类评分:通用文本 40.6/100 (n=18);速度与延迟 39.3/100 (n=6)。 评分方法与数据来源
近期走势
口碑在如何变化
近 30 天,每点代表截至该日采样时刻的近 7 天滚动体感分。纵轴随数据调整;缺失或样本不足处断开,今天仍在更新。
点按或用 ← → 查看日期、分数与样本量;空白日期暂无足够数据。
查看每日读数与样本量
| 日期 | 体感分 | n | 评分窗口(UTC) |
|---|---|---|---|
| 2026-10-03 | 30.7 | 8 | 2026-09-26T13:42:04+00:00 – 2026-10-03T13:42:04+00:00 |
| 2026-10-02 | 30.7 | 8 | 2026-09-25T23:57:03+00:00 – 2026-10-02T23:57:03+00:00 |
| 2026-10-01 | 30.7 | 8 | 2026-09-24T23:57:02+00:00 – 2026-10-01T23:57:02+00:00 |
| 2026-09-30 | 30.7 | 8 | 2026-09-23T23:57:03+00:00 – 2026-09-30T23:57:03+00:00 |
| 2026-09-29 | 31.7 | 9 | 2026-09-22T23:57:03+00:00 – 2026-09-29T23:57:03+00:00 |
| 2026-09-28 | 31.8 | 9 | 2026-09-21T23:57:03+00:00 – 2026-09-28T23:57:03+00:00 |
| 2026-09-27 | — | 3 | 2026-09-20T23:57:03+00:00 – 2026-09-27T23:57:03+00:00 |
| 2026-09-26 | — | 2 | 2026-09-19T23:57:03+00:00 – 2026-09-26T23:57:03+00:00 |
| 2026-09-25 | — | 2 | 2026-09-18T23:57:03+00:00 – 2026-09-25T23:57:03+00:00 |
| 2026-09-24 | — | 2 | 2026-09-17T23:57:03+00:00 – 2026-09-24T23:57:03+00:00 |
| 2026-09-23 | — | 1 | 2026-09-16T23:57:03+00:00 – 2026-09-23T23:57:03+00:00 |
| 2026-09-22 | — | 2 | 2026-09-15T23:57:03+00:00 – 2026-09-22T23:57:03+00:00 |
| 2026-09-21 | — | 2 | 2026-09-14T23:57:03+00:00 – 2026-09-21T23:57:03+00:00 |
| 2026-09-20 | — | 2 | 2026-09-13T23:57:03+00:00 – 2026-09-20T23:57:03+00:00 |
| 2026-09-19 | — | 2 | 2026-09-12T23:57:03+00:00 – 2026-09-19T23:57:03+00:00 |
| 2026-09-18 | — | 2 | 2026-09-11T23:57:02+00:00 – 2026-09-18T23:57:02+00:00 |
| 2026-09-17 | — | 3 | 2026-09-10T23:57:03+00:00 – 2026-09-17T23:57:03+00:00 |
| 2026-09-16 | — | 3 | 2026-09-09T23:57:03+00:00 – 2026-09-16T23:57:03+00:00 |
| 2026-09-15 | — | 2 | 2026-09-08T23:57:03+00:00 – 2026-09-15T23:57:03+00:00 |
| 2026-09-14 | — | 1 | 2026-09-07T23:57:03+00:00 – 2026-09-14T23:57:03+00:00 |
| 2026-09-13 | — | 1 | 2026-09-06T23:57:03+00:00 – 2026-09-13T23:57:03+00:00 |
| 2026-09-12 | — | 3 | 2026-09-05T23:57:03+00:00 – 2026-09-12T23:57:03+00:00 |
| 2026-09-11 | — | 3 | 2026-09-04T23:00:04+00:00 – 2026-09-11T23:00:04+00:00 |
| 2026-09-10 | — | 4 | 2026-09-03T23:00:05+00:00 – 2026-09-10T23:00:05+00:00 |
| 2026-09-09 | 58.5 | 6 | 2026-09-02T23:00:04+00:00 – 2026-09-09T23:00:04+00:00 |
| 2026-09-08 | 50.7 | 8 | 2026-09-01T23:00:03+00:00 – 2026-09-08T23:00:03+00:00 |
| 2026-09-07 | 50.7 | 8 | 2026-08-31T23:00:03+00:00 – 2026-09-07T23:00:03+00:00 |
| 2026-09-06 | 46.9 | 7 | 2026-08-30T23:00:02+00:00 – 2026-09-06T23:00:02+00:00 |
| 2026-09-05 | 42.6 | 6 | 2026-08-29T23:00:03+00:00 – 2026-09-05T23:00:03+00:00 |
| 2026-09-04 | 38.4 | 5 | 2026-08-28T23:00:03+00:00 – 2026-09-04T23:00:03+00:00 |
Hacker News · 30.7 · 8 条评分评论 · 查看该平台评价 →
不同社区,同一模型
各平台怎么看
各平台独立计算,样本量和讨论人群不同。平台分数不取简单平均;评分评论不包含仅讨论额度的评论。
分数范围 0–100。点击平台查看趋势及对应评价。最新评论时间不代表爬虫运行状态。
查看长期变化与分析
此版本尚未纳入发布后长期追踪;上方为近7天口碑。
社区评论
精选近 7 天的不同意见,摘录条数不代表真实好评比例。
2 条精选摘录
展开原文
mostly sonnet 5 or gemini flash 3.8
查看评论上下文
6.1 Sol又慢又懒,沟通能力差,代码质量不如其他模型
机器翻译展开原文
6.1 Sol is slow, lazy, does not communicate well and code quality is not on par with other models
展开原文
it was so much better than opus or sonnet 5
展开原文
sonnet 5 was basically unusable
展开原文
both Sonnet or Fable would have nailed every time
查看评论上下文
但在几个任务之后,我意识到它遗漏了复杂任务的大量细节,而Fable绝不会遗漏这些。它反而创建了一个能通过自己测试的版本,并断言工作已完成(公平地说,所有LLM都会这样做)。但与5.1相比,它就像是在故意走捷径来让工作"完成"。看起来它被调校成会修剪任务中"非必要"的部分,以便更快地到达终点。 而且我不认为这仅限于那些有解释空间的大型任务。我发现在小得多的附带任务上它也会这样做,而Sonnet或Fable每次都能完美完成这些任务。
机器翻译展开原文
But after a few tasks I realised it was omitting huge details of complex tasks that Fable just would not have missed. It instead creates a version that passes its own tests and asserts that the job is done (which to be fair, all LLMs do). But in comparison to 5.1, its like it takes intentional shortcuts to get the job 'done'. It looks like it's been tuned to trim "non-essential" aspects of a task to get to the end faster. And I don't think this is isolated to large tasks that have room for interpretation. I've caught doing the same with much smaller side tasks that both Sonnet or Fable would have nailed every time.
展开原文
Sonnet 5 was so horrible
展开原文
Sonnet 5 to perform worse across all of our evals, especially against time
展开原文
I would never get an answer back even for very simple prompts. It would just churn on nothing and return max token usage reached
展开原文
Sonnet 5 is genuinely a terrible model for me
展开原文
how obnoxious and lazy Sonnet 5 was
展开原文
It does worse than Sonnet 5. Mainly because it is more reluctant to keep going to get an answer
查看评论上下文
我构建了一种对抗性的深奥编程语言来对LLM模型进行基准测试,刚在Sonnet 5.5上运行了它。它的表现比Sonnet 5还差。
机器翻译展开原文
I built an adversarial esoteric programming language to benchmark LLM models and just ran it on Sonnet 5.5 It does worse than Sonnet 5.
展开原文
I even replaced it with sonnet 5 recently because I wasn't very impressed with the results
查看评论上下文
在工作中我们有 copilot,即使有 750 美元的额度,我也很少使用 opus,我最近甚至用 sonnet 5 替换了它,因为我对结果不太满意,等不及下周我们启用 opus 5.5 了。
机器翻译展开原文
At work we have copilot and even with 750$ credits I use opus very sparingly and I even replaced it with sonnet 5 recently because I wasn't very impressed with the results, can't wait we enable opus 5.5 next week.
近 7 天暂时没有符合此筛选条件的精选评论。