Claude · Last 7 days
Claude Sonnet 5.5
0–100 · higher means more positive community experience, not benchmark performance.
Updated 2026-10-02 15:02 UTC · 2026-09-25 – 2026-10-02 UTCWhat people think
One user added a mini router that dispatches harder coding slices to Sonnet 5.5, and the setup improved first-pass implementation quality. One user, after briefly reviewing Claude Sonnet 5.5's Thomson problem proof, sarcastically said anyone who finds inspiration from it should just become a computer.
About the sample and scores
As of 2026-10-02 15:02:03 UTC, Claude Sonnet 5.5 in the Claude family has 122 explicitly attributed community comments in the last 7 days; its community score is 66.2/100. Available category scores include General text 66.2/100 (n=59); Coding 76.4/100 (n=9); Reasoning 66.2/100 (n=26). Scoring method and sources
RECENT EXPERIENCE
How the Experience Index is changing
30 days of rolling 7-day scores at each daily snapshot. The vertical scale adapts to the data. Missing or insufficient samples leave gaps; today is still updating.
Tap or use ← → for dates, scores and sample sizes. Blank dates have insufficient data.
Daily readings & sample sizes
| Date | Score | n | Scoring window (UTC) |
|---|---|---|---|
| 2026-10-02 | 60.9 | 51 | 2026-09-25T15:02:03+00:00 – 2026-10-02T15:02:03+00:00 |
| 2026-10-01 | 60.4 | 50 | 2026-09-24T23:57:02+00:00 – 2026-10-01T23:57:02+00:00 |
| 2026-09-30 | 59.9 | 47 | 2026-09-23T23:57:03+00:00 – 2026-09-30T23:57:03+00:00 |
| 2026-09-29 | 62.5 | 39 | 2026-09-22T23:57:03+00:00 – 2026-09-29T23:57:03+00:00 |
Zhihu · 60.9 · 51 scored comments · Read platform reviews →
ONE MODEL, DIFFERENT COMMUNITIES
Across the communities
Scored separately for each platform, with different samples and audiences. Platform scores are not simply averaged; scored counts exclude quota-only comments.
Scores range from 0 to 100. Select a platform for its trend and reviews. Latest opinion time is not crawler health.
Community reviews
Selected comments from the last 7 days. The balance of excerpts does not represent the share of positive reviews.
15 selected excerpts
Community reports about limits cannot establish official allowances.
Show original
能在这里面产生灵感的可以当计算机去了
Comment context
Some responses claim this is Anthropic's excessive promotion, but I feel I need to stand up for Anthropic. The reason is simple—this is not Anthropic's release at all... The original message was from a company doing various LLM benchmarks called Vals.AI. Here is the original text: We asked ten Claude Sonnet 5.5 agents to use Lean to prove the lowest-energy arrangement of seven electrons on a sphere (the Thomson problem, with N=7). Within 15 hours, they produced a 17,895-line proof, accepted by the Lean kernel, showing that the answer is a pentagonal bipyramid.
Machine translatedShow original
有些回答说这是anthropic的过度宣传,我觉得我得为anthropic站一下台。理由很简单,这根本不是anthropic发的…… 原消息是一家做各种LLM Benchmark的公司Vals.AI发的,这是原文: We asked ten Claude Sonnet 5.5 agents to use Lean to prove the lowest-energy arrangement of seven electrons on a sphere (the Thomson problem, with N=7). Within 15 hours, they produced a 17,895-line proof, accepted by the Lean kernel, showing that the answer is a pentagonal bipyramid.
Comment context
I learned through my testing that sonnet 5.5 sometimes spawns opus 5.5 subagents. Might be one of the main reasons why sonnet 5.5 feels a lot more expensive for people.
And, this is not about the API price, but the percentage weekly usage of the subscription.
Comment context
I actually added a mini router in front of it this week, and the harder slices get dispatched to sonnet 5.5 instead, pretty effective so far, it improved first-pass implementation quality.
Show original
这玩意儿哪里都不如Opus 5.5,也是给A畜学到国产小参数做题家模型雷霆大思考刷分的精髓了
Comment context
So I don't know what's your process but I would suggest models wise to just use Opus 5.5 high or medium as the orchestrator to supervise a Sonnet 5.5 high implementer and use 6.1 sol xhigh (or medium if you notice xhigh is not sustainable for you) as a reviewer.
Comment context
It really depends on the task, but it's often either sonnet 5.5 high or opus 5.5 med.
Show original
丫抠字儿太快了 找她吹牛逼容易跟不上思路
Show original
而且很快
Show original
说人话的,废话也少了
Show original
依靠无限制拉长思考来换取所谓的更高智商,感觉可能不是正确方向
Show original
sonnet5.5审美确实一眼高级
Show original
追问,发现有幻觉
Comment context
I discussed a JEV-related question and found that Company A's models now have a certain characteristic.
Machine translatedShow original
我拿一个跟JEV相关的问题讨论了一下,发现A家的模型现在有个特点。
Show original
在智能体命令行编程测试 Terminal-Bench 4.0 中,它以 70.6% 的高分拿下 全球第一 ,相比 Sonnet 5 的 10.3% 属于跨越式暴涨
Show original
A社的账号"娇贵"啊,隔三岔五给你大陆ip全干挺了
Show original
才发现Sonnet5.5还挺耐用啊
Show original
思考用更多token导致更慢
Show original
免费用户也能用几下sonnet5.5,反观隔壁免费用户只能用luna
Show original
同思考档位的sonnet 5.5 比 opus 消耗token更多,效率更低
Show original
单价只有Opus一半,额度还给得更多,还学会了O社时不时发重置次数
Comment context
I've never once hit a safeguard refusal. Dunno what you guys are doing
Comment context
If Astra is that great, how come its code introduces so many bugs? Give Astra a complicated task on Ultra, and then after give the prompt "review" - you will get about 5 problems.
Comment context
Tokens | cost to serve/run the same agentic task | performance across multiple benchmarks
No selected comments for this filter in the last 7 days.