Claude · 近 7 天
Claude Opus 5
0–100 分,越高代表社区使用体验越正面,不代表能力测试成绩。
更新于 2026-09-09 16:10 UTC · 2026-09-02 – 2026-09-09 UTC社区口碑分32.7/100184 条评论 · 近 7 天
截至 2026-09-09 16:00:04 UTC,Claude 家族的 Claude Opus 5 在近 7 天有 184 条明确归属该版本的社区反馈,社区口碑分为 32.7/100。部分有效分类评分:通用文本 16.4/100 (n=96);编程 52.9/100 (n=30);推理 56.2/100 (n=34)。 评分方法与数据来源
查看长期变化与分析
模型生命周期
Claude Opus 5
发布于 2026-07-24 · 返回 Claude
精确版本信号覆盖 2026-08-07–2026-09-08 · n=622
统计窗口:2026-08-12 00:00:00 至 2026-09-09 00:00:00 UTC · 评分方法:experience_score_v2。
趋势历史仍在收集中
当前窗口已达到要求;用于判定的独立分段已有 2/4 个。
生命周期每日更新 · 最新完整 UTC 日期 2026-09-08
- 固定 14 天基准
- 45.5 n=37
- 最近 28 天
- 28.2 n=594
- 较基准变化
- −17.3
- 90 天斜率
- – 分 / 30 天
14 天滚动体感趋势
每个每日更新的点汇总此前 14 个完整 UTC 日。虚线是固定发布基准;正式趋势判决仍使用互不重叠的独立 14 天分段。
用于趋势判定的独立 14 天分段已就绪 2/4 个。
哪些维度解释了变化
分别展示各维度内部的口碑变化与讨论构成变化。这是观察性贡献,不是因果证明。
主要负向口碑贡献:通用文本。
| 维度 | 基准 → 当前 | 维度变化 | 权重占比 | 口碑变化贡献 | 讨论构成贡献 |
|---|---|---|---|---|---|
| 通用文本 | 基准 → 当前36.5 → 16.4n=18 → 285 | 维度变化−20.1 | 权重占比44.9% → 46.8% | 口碑变化贡献−9.23 | 讨论构成贡献−0.20 |
| 编程 | 基准 → 当前样本不足n=3 → 101 | 维度变化– | 权重占比– | 口碑变化贡献– | 讨论构成贡献– |
| 图像/视觉 | 基准 → 当前样本不足n=0 → 0 | 维度变化– | 权重占比– | 口碑变化贡献– | 讨论构成贡献– |
| 本地部署 | 基准 → 当前样本不足n=0 → 0 | 维度变化– | 权重占比– | 口碑变化贡献– | 讨论构成贡献– |
| 推理 | 基准 → 当前样本不足n=9 → 126 | 维度变化– | 权重占比– | 口碑变化贡献– | 讨论构成贡献– |
| 角色扮演/创作 | 基准 → 当前样本不足n=0 → 9 | 维度变化– | 权重占比– | 口碑变化贡献– | 讨论构成贡献– |
| 安全与拒答 | 基准 → 当前样本不足n=2 → 14 | 维度变化– | 权重占比– | 口碑变化贡献– | 讨论构成贡献– |
| 速度与延迟 | 基准 → 当前样本不足n=5 → 59 | 维度变化– | 权重占比– | 口碑变化贡献– | 讨论构成贡献– |
| 视频生成 | 基准 → 当前样本不足n=0 → 0 | 维度变化– | 权重占比– | 口碑变化贡献– | 讨论构成贡献– |
样本不足维度的未解释贡献:−7.88 分。
生命周期分沿用主指数的体验信号权重,n≥30 后不做样本收缩。它衡量公开用户感知,不代表模型能力,也不能证明后端原因。
社区评论
精选近 7 天的不同意见,摘录条数不代表真实好评比例。
2 条精选摘录
展开原文
opus 5 is great at coding, as long as you use hooks to make him shut the fuck up and to block commentary as well since he tends to write 3 comments for every 3 lines of code
展开原文
I had good results with Opus 5 too in spite of its painful tone & verbosity. Very competent model.
展开原文
5.6 Sol/Opus 5 and better models can for example handle my governments ridiculous electronic forms and all associated bureaucracy much better than I can
展开原文
most convoluted and jargon-filled output
展开原文
In cases like Opus 5 it's near impossible to overcome
展开原文
Opus diligently verifies their claims and fixes things
展开原文
"ah, but Fable is sooooooo dangerous, as is Opus 5" is just a terrible hand wave
展开原文
I will be a Opus 5 slave for the foreseeable future
展开原文
Opus 5 is my best coder still, but Astra is finally producing comparable results to it. They both burn comparable tokens on a clean context spin-up, but Opus burns less once I'm iterating on an existing context.
展开原文
Opus 5 tends give too much priority to minor typos/linguistic choices, adopts a pedantic and adversarial approach to doc reviews, proposes unwise strategies and glosses over important details. It also has a somewhat shaky understanding of E
展开原文
Opus 5 as subagents
展开原文
100% agree with your description of Opus 5 lol
展开原文
Opus 5 and Sonnet 5 the experiences people have are very inconsistent. For some people Opus 5 is solid, others like me do not trust it outside of mockups and plan mode
展开原文
Opus 5 on High is orders of magnitude slower burn - but I do see a difference in result quality
展开原文
Same experience here. Usage drain too fast as well which wasn't the case before or few days ago.
展开原文
opus 5 is better than both but you need to make it recheck everything it made with cold reviewer because it always makes mistake
近 7 天暂时没有符合此筛选条件的精选评论。