Claude · 近 7 天
Claude Opus 5.5
0–100 分,越高代表社区使用体验越正面,不代表能力测试成绩。
更新于 2026-10-04 08:32 UTC · 2026-09-27 – 2026-10-04 UTC讨论概览
主要讨论什么
全部平台 · 近7天 784 条合格评论,包含额度反馈;按来源与评论去重。
- 通用文本 · 306 条:正面 201 / 负面 70 / 褒贬或中性 35
- 推理 · 211 条:正面 185 / 负面 17 / 褒贬或中性 9
- 速度与延迟 · 109 条:正面 79 / 负面 19 / 褒贬或中性 11
- 使用额度 · 97 条:正面 52 / 负面 26 / 褒贬或中性 19
同一评论可涉及多个维度,各项不可相加;方向计数不等于加权评分。链接展示该维度的精选原文。
近期变化
近7天滚动体感分较7天前 −6.8 分。讨论构成也会影响分数,这不是原因诊断。
样本分布与讨论集中度
Hacker News 88 · Reddit 625 · v2ex 1 · 小红书 10 · 知乎 60
可识别 324 个讨论,覆盖 696/784 条评论;最大单帖 28 条。其余讨论归属未知,不推算独立用户数。
个体观点选摘
有用户反馈 Sol 6.1 的每秒输出 token 数低于 Opus 5.5,导致 Sol 6.1 完成任务耗时极长。 有用户表示 Claude Opus 5.5 使用起来太贵了。
样本与评分说明
截至 2026-10-04 08:32:03 UTC,Claude 家族的 Claude Opus 5.5 在近 7 天有 784 条明确归属该版本的社区反馈,社区口碑分为 77.3/100。部分有效分类评分:通用文本 72.5/100 (n=306);编程 82.1/100 (n=78);推理 89.1/100 (n=211)。 评分方法与数据来源
近期走势
口碑在如何变化
近 30 天,每点代表截至该日采样时刻的近 7 天滚动体感分。纵轴随数据调整;缺失或样本不足处断开,今天仍在更新。
点按或用 ← → 查看日期、分数与样本量;空白日期暂无足够数据。
查看每日读数与样本量
| 日期 | 体感分 | n | 评分窗口(UTC) |
|---|---|---|---|
| 2026-10-04 | 81.9 | 59 | 2026-09-27T08:32:03+00:00 – 2026-10-04T08:32:03+00:00 |
| 2026-10-03 | 80.0 | 59 | 2026-09-26T23:57:02+00:00 – 2026-10-03T23:57:02+00:00 |
| 2026-10-02 | 80.1 | 60 | 2026-09-25T23:57:03+00:00 – 2026-10-02T23:57:03+00:00 |
| 2026-10-01 | 81.6 | 59 | 2026-09-24T23:57:02+00:00 – 2026-10-01T23:57:02+00:00 |
| 2026-09-30 | 84.1 | 58 | 2026-09-23T23:57:03+00:00 – 2026-09-30T23:57:03+00:00 |
| 2026-09-29 | 83.8 | 64 | 2026-09-22T23:57:03+00:00 – 2026-09-29T23:57:03+00:00 |
| 2026-09-28 | 84.2 | 65 | 2026-09-21T23:57:03+00:00 – 2026-09-28T23:57:03+00:00 |
| 2026-09-27 | 82.7 | 41 | 2026-09-20T23:57:03+00:00 – 2026-09-27T23:57:03+00:00 |
| 2026-09-26 | 84.5 | 34 | 2026-09-19T23:57:03+00:00 – 2026-09-26T23:57:03+00:00 |
| 2026-09-25 | 82.9 | 30 | 2026-09-18T23:57:03+00:00 – 2026-09-25T23:57:03+00:00 |
| 2026-09-24 | 81.0 | 26 | 2026-09-17T23:57:03+00:00 – 2026-09-24T23:57:03+00:00 |
| 2026-09-23 | 76.2 | 19 | 2026-09-16T23:57:03+00:00 – 2026-09-23T23:57:03+00:00 |
知乎 · 81.9 · 59 条评分评论 · 查看该平台评价 →
不同社区,同一模型
各平台怎么看
各平台独立计算,样本量和讨论人群不同。平台分数不取简单平均;评分评论不包含仅讨论额度的评论。
分数范围 0–100。点击平台查看趋势及对应评价。最新评论时间不代表爬虫运行状态。
查看长期变化与分析
此版本尚未纳入发布后长期追踪;上方为近7天口碑。
社区评论
精选近 7 天的不同意见,摘录条数不代表真实好评比例。
2 条精选摘录
额度评价来自社区体验,不能据此推算官方实际配额。
展开原文
TPS is lower on Sol 6.1 than Opus 5.5. To the point that 6.1 takes forever to finish a task.
展开原文
Imagine you had a burger that usually tasted good but they stopped giving you free condiments. Then across the street is a burger with free condiments, fries, and a drink for the same price.
展开原文
Opus 5.5 is as capable as Astra and in some cases better than it and I can run it on their stupid $20 plan
展开原文
too expensive to use
展开原文
The shittiest version of Opus 5.5 will always be better than Sol 6.1.
展开原文
Opus has spent 2 hours fixing what Astra spent 4 hours doing this morning. The quality is night and day between the two
展开原文
I enjoy my Opus 5.5
展开原文
great and cheaper if you compare to fable with close quality in coding
展开原文
The only thing I can tell is the sol is slower, but def outputs the same work as opus.
展开原文
they can't even answer if robots replace all human for work then who tf is going to pay for the robots by whose money
展开原文
the 5 hour limit basically doesn’t exist, it literally drops 1-2% over that same period
展开原文
sure 1/1 token usage might be better with 5.5 but I am pretty sure 4.8 was using less x<---/x than 5.5 has ben using
展开原文
Getting more usage out of Opus 5.5
展开原文
Opus 5.5 got 1 full project and 12 other side-projects far enough that I'm gonna be shipping them tomorrow, all in just a single week
展开原文
I used Opus 5.5 with a 2m LOC codebase to refactor the versioning and publishing aspect of the system.
展开原文
125 s against 75 s
查看评论上下文
在更改计划前测量了我的消耗:将 verify 命令作为第二轮运行的费用是每个任务 $0.96,而 Opus 5.5 普通运行是 $0.39,用时 125 秒对比 75 秒。
机器翻译展开原文
Measured my own burn before changing plans: running the verify command as a second turn cost $0.96 a task against $0.39 plain on Opus 5.5, 125 s against 75 s.
展开原文
the usage rate is superior to Opus 5.5
查看评论上下文
我换到了Sonnet 5.5,它需要多照顾一点,但使用率优于Opus 5.5。
机器翻译展开原文
I switched over to Sonnet 5.5 and it needs a bit more hand holding, but the usage rate is superior to Opus 5.5.
展开原文
it is muuuch more likely to be confidently incorrect
展开原文
Anthropic needs to figure something out with their restrictions.
展开原文
Claude literally has no idea what he writes, literally feels like talking to a clueless person
展开原文
Yeah it's just their security protocols.
查看评论上下文
GPT 6.1和Opus 5.5都拒绝输入smoke test的密码是怎么回事?我为我的测试和staging环境设置了一个本地环境,我向Claude和GPT都解释了这都是本地的和staging环境,但两个模型都拒绝为我想要它们走一遍并创建测试账户编造密码。它们都一直停下来让我自己来做。
机器翻译展开原文
What’s with GPT 6.1 and Opus 5.5 both refusing to type passwords for smoke tests? I’ve got a local environment setup for my testing and staging environment, I explain to both Claude and GPT it’s all local and a staging environment and both models refuse to make up passwords for test accounts I want them to walk through and create. They both keep stopping and asking me to do it.
展开原文
takes about twice as long and twice as many tokens as 5.5
查看评论上下文
例如,如果你让Opus 5处理一堆来自审查代码的子代理扩散的审查,然后将它们合成一份最终报告。与5.5相比,它能保留更多的实际缺陷,但差距不大。但它也会保留更多的误报,而且耗时大约是5.5的两倍,消耗的token也大约是5.5的两倍。
机器翻译展开原文
For example if you have Opus 5 process a bunch of reviews from an subagent fanout that is reviewing code, then synthesizing those into a final report. It keeps more of the actual defects vs 5.5, by a small margin. But it also keeps more false positive and takes about twice as long and twice as many tokens as 5.5
展开原文
Between cost, spend, and code quality I can literally prove that something has changed.
查看评论上下文
在过去的六天里,Opus 5.5(O5.5)在 Claude Code 中对我的表现非常出色。开箱即用的架构优先、DRY/SOLID 编码,出色的沟通风格,惊人的 token 效率。从上周三开始,我基本上不间断地使用它,每天大约 12 个小时。 但就在今晚我的月度限额重置前后,我注意到沟通风格和编码行为发生了*极端*的转变,感觉可疑地像 Opus 5(O5)。在过去一周里,O5.5 会接受我的需求并立即开始工作,通常几分钟内就能完成一个功能,在实现、用户体验、架构和 token 效率方面都做得令人难以置信。而今晚,我看到的情况截然不同。
机器翻译展开原文
Opus 5.5 (O5.5) has been exceptional for me in Claude Code over the last six days. Architecture-first, DRY/SOLID coding out of the box, exceptional communication style, phenomenal token efficiency. I’ve been working with it basically nonstop, ~12 hours a day, since last Wednesday. But right around the time my monthly limit reset this evening, I noticed an extreme shift in communication style and coding behavior that feels suspiciously like Opus 5 (O5). For the last week, O5.5 would take my requirements and immediately get to work, usually knocking out a feature in minutes and doing an incredible job across implementation, UX, architecture, and token efficiency. Tonight, I’m seeing something very different.
O5.5 一直始终如一地尊重 DRY 原则和现有架构,而今晚创建的模块却突然乐于重新发明之前已有的每一个轮子。能力上的差异并不微妙。 *另一个*迹象是 token 使用量。与 O5 相比,O5.5 的效率惊人,而 O5 绝对是一个吞噬 token 的怪物。我用 O5.5 以一小部分 token 使用量完成了多得多的功能工作。而今晚,简单的任务和相对较小的功能突然以更接近 O5 的速度吞噬 token。 我对此特别敏感,因为每次达到消费上限时我都必须申请提高额度。用 O5 时,我几乎每天都要提交一次请求。用 O5.5 时,我整个星期都不需要提交一次。然后今晚,我在大约一个小时内从约 70% 的使用量跳到了 90%。
机器翻译展开原文
Where O5.5 consistently respected DRY principles and the existing architecture, the modules created tonight are suddenly happy to reinvent every wheel that came before them. The difference in ability is not subtle. The other tell is token usage. O5.5 has been startlingly efficient compared with O5, which was an absolute token-eating monster. I’ve been getting substantially more feature work done with O5.5 at a fraction of the token usage. Tonight, simple tasks and relatively small features are suddenly devouring tokens at a pace that feels much more like O5. I’m especially sensitive to this because I have to request an increase in my spend limit every time I hit it. On O5, I was sending a request almost every day. On O5.5, I hadn’t needed to send a single one all week. Then tonight I jumped from roughly 70% to 90% usage in about an hour.
展开原文
in terms of creativity, creating new things, yeah opus 5.5 is fantastic
展开原文
it catches many mistakes made by Opus 5.5
查看评论上下文
SOL 5.6 和 6.1 都是高水平后端编码的怪兽。
机器翻译展开原文
SOL 5.6 and 6.1 both are monsters for high level backend coding.
展开原文
Opus 5.5 (I use it for work writing C#) seems faster than Opus 5, but it doesn't seem demonstrably better to my eyes and is still prone to word vomit
展开原文
I have yet to have Opus 5.5 block me from doing any of my usual security testing, validation, and remediation work under the CVP.
展开原文
Opus is clearly better in logic and processing steps, but often skips some or does something else entirely
展开原文
graphics would be 10x better
查看评论上下文
它始于我输入到Claude Code中的一条提示词:“*仅使用像素直接重制一个精美的宝可梦红重制版……绝对的像素艺术完美。
机器翻译展开原文
It started from one prompt I typed into Claude Code: "*using only pixels directly remake a beautiful pokemon red remake … absolute pixel art perfection.
* 完全没有图像文件。 游戏逐像素绘制到 320×180 的屏幕上。
机器翻译展开原文
* No image files at all. The game draws into a 320×180 screen pixel by pixel.
展开原文
Opus 5.5 has significantly better vision, which might mean for game dev you're using it more
近 7 天暂时没有符合此筛选条件的精选评论。