Claude · 近 7 天
Claude Sonnet 5.5
0–100 分,越高代表社区使用体验越正面,不代表能力测试成绩。
更新于 2026-10-03 08:42 UTC · 2026-09-26 – 2026-10-03 UTC讨论概览
主要讨论什么
全部平台 · 近7天 129 条合格评论,包含额度反馈;按来源与评论去重。
- 通用文本 · 62 条:正面 43 / 负面 18 / 褒贬或中性 1
- 推理 · 29 条:正面 20 / 负面 7 / 褒贬或中性 2
- 速度与延迟 · 25 条:正面 18 / 负面 6 / 褒贬或中性 1
- 使用额度 · 11 条:正面 3 / 负面 6 / 褒贬或中性 2
同一评论可涉及多个维度,各项不可相加;方向计数不等于加权评分。链接展示该维度的精选原文。
近期变化
近期变化:历史不足,暂不判断改善或下降。
样本分布与讨论集中度
Hacker News 4 · Reddit 66 · 知乎 59
可识别 44 个讨论,覆盖 125/129 条评论;最大单帖 51 条。其余讨论归属未知,不推算独立用户数。
个体观点选摘
有用户建议在购买100美元套餐前应更充分地使用Sonnet 5.5,大多数人过早切到Opus 了。 有用户稍看了下Claude Sonnet 5.5生成的汤姆逊问题证明后,嘲讽地说能从这种证明中获得灵感的人不如直接去当计算机。
样本与评分说明
截至 2026-10-03 08:42:04 UTC,Claude 家族的 Claude Sonnet 5.5 在近 7 天有 129 条明确归属该版本的社区反馈,社区口碑分为 66.7/100。部分有效分类评分:通用文本 66.0/100 (n=62);编程 76.4/100 (n=9);推理 68.0/100 (n=29)。 评分方法与数据来源
近期走势
口碑在如何变化
近 30 天,每点代表截至该日采样时刻的近 7 天滚动体感分。纵轴随数据调整;缺失或样本不足处断开,今天仍在更新。
点按或用 ← → 查看日期、分数与样本量;空白日期暂无足够数据。
查看每日读数与样本量
| 日期 | 体感分 | n | 评分窗口(UTC) |
|---|---|---|---|
| 2026-10-03 | 68.3 | 61 | 2026-09-26T08:42:04+00:00 – 2026-10-03T08:42:04+00:00 |
| 2026-10-02 | 67.7 | 60 | 2026-09-25T23:57:03+00:00 – 2026-10-02T23:57:03+00:00 |
| 2026-10-01 | 66.3 | 57 | 2026-09-24T23:57:02+00:00 – 2026-10-01T23:57:02+00:00 |
| 2026-09-30 | 66.9 | 48 | 2026-09-23T23:57:03+00:00 – 2026-09-30T23:57:03+00:00 |
| 2026-09-29 | 63.8 | 41 | 2026-09-22T23:57:03+00:00 – 2026-09-29T23:57:03+00:00 |
Reddit · 68.3 · 61 条评分评论 · 查看该平台评价 →
不同社区,同一模型
各平台怎么看
各平台独立计算,样本量和讨论人群不同。平台分数不取简单平均;评分评论不包含仅讨论额度的评论。
分数范围 0–100。点击平台查看趋势及对应评价。最新评论时间不代表爬虫运行状态。
查看长期变化与分析
此版本尚未纳入发布后长期追踪;上方为近7天口碑。
社区评论
精选近 7 天的不同意见,摘录条数不代表真实好评比例。
14 条精选摘录
额度评价来自社区体验,不能据此推算官方实际配额。
展开原文
the usage rate is superior to Opus 5.5
查看评论上下文
说实话,我觉得这就是个代餐。因为opus 5.5过于优秀,无论是价格还是能力,我自己用量也没大到那种程度,所以我现在的工作全部都交给opus 5.5也很香,也就没必要再关注sonnet了。
展开原文
i'd push sonnet harder before paying for 100$
展开原文
Actually now with sonnet 5.5 Claude's $20 plan would be pretty fantastic too
查看评论上下文
有些回答说这是anthropic的过度宣传,我觉得我得为anthropic站一下台。理由很简单,这根本不是anthropic发的…… 原消息是一家做各种LLM Benchmark的公司Vals.AI发的,这是原文: We asked ten Claude Sonnet 5.5 agents to use Lean to prove the lowest-energy arrangement of seven electrons on a sphere (the Thomson problem, with N=7). Within 15 hours, they produced a 17,895-line proof, accepted by the Lean kernel, showing that the answer is a pentagonal bipyramid.
展开原文
the way it talks is so weird and unclear compared to 5.0
展开原文
sonnet 5.5 feels a lot more expensive for people
查看评论上下文
我通过测试了解到,sonnet 5.5有时会生成opus 5.5子代理。这可能是人们觉得sonnet 5.5贵很多的主要原因之一。
机器翻译展开原文
I learned through my testing that sonnet 5.5 sometimes spawns opus 5.5 subagents. Might be one of the main reasons why sonnet 5.5 feels a lot more expensive for people.
而且,这不是关于API价格的问题,而是订阅的每周使用百分比。
机器翻译展开原文
And, this is not about the API price, but the percentage weekly usage of the subscription.
展开原文
improved first-pass implementation quality
查看评论上下文
我这周实际上在它前面加了一个迷你路由器,更难的切片会被分派给sonnet 5.5,到目前为止相当有效,它提高了首次实现的代码质量。
机器翻译展开原文
I actually added a mini router in front of it this week, and the harder slices get dispatched to sonnet 5.5 instead, pretty effective so far, it improved first-pass implementation quality.
展开原文
I wouldn't trust sonnet to do give me any architecture/design advices.
展开原文
often either sonnet 5.5 high
查看评论上下文
真的取决于任务,但它通常要么是 sonnet 5.5 高级要么是 opus 5.5 中等。
机器翻译展开原文
It really depends on the task, but it's often either sonnet 5.5 high or opus 5.5 med.
查看评论上下文
我拿一个跟JEV相关的问题讨论了一下,发现A家的模型现在有个特点。
展开原文
false positive guardrail blocks
展开原文
I asked it to write gemini on doctor strange. It wouldnt.
查看评论上下文
我一次都没遇到过安全机制拒绝。不知道你们在搞什么
机器翻译展开原文
I've never once hit a safeguard refusal. Dunno what you guys are doing
展开原文
I've never once hit a safeguard refusal
展开原文
it can be amazing
展开原文
What the OP stated happens with Fable 5.1, Opus 5.5, Sonnet 5.5 too.
查看评论上下文
如果 Astra 那么厉害,为什么它的代码会引入这么多 bug? 给 Astra 在 Ultra 上一个复杂任务,然后之后输入提示词"review"——你会得到大约 5 个问题。
机器翻译展开原文
If Astra is that great, how come its code introduces so many bugs? Give Astra a complicated task on Ultra, and then after give the prompt "review" - you will get about 5 problems.
展开原文
Sonnet 5.5...: 382,000 ($463) (94.6%)
查看评论上下文
Tokens | 服务/运行相同代理任务的成本 | 多个基准测试中的性能
机器翻译展开原文
Tokens | cost to serve/run the same agentic task | performance across multiple benchmarks
近 7 天暂时没有符合此筛选条件的精选评论。