Claude · Last 7 days
Claude Opus 5
0–100 · higher means more positive community experience, not benchmark performance.
Updated 2026-10-02 22:02 UTC · 2026-09-25 – 2026-10-02 UTCDiscussion overview
What the full sample discusses
All platforms · 75 eligible comments over 7 days, including usage-limit feedback; deduplicated by source and comment.
- General text · 46 comments: 5 positive / 37 negative / 4 mixed or neutral
- Reasoning · 16 comments: 5 positive / 11 negative / 0 mixed or neutral
- Speed & latency · 6 comments: 0 positive / 5 negative / 1 mixed or neutral
- Usage limits · 5 comments: 1 positive / 3 negative / 1 mixed or neutral
Comments may cover multiple dimensions; counts are not additive or equivalent to the weighted score. Links show selected source excerpts.
Recent change
Rolling 7-day score: −13.4 points vs 7 days ago. Discussion mix can affect scores; this does not establish a cause.
Source coverage and discussion concentration
Hacker News 6 · Reddit 63 · RedNote 1 · Zhihu 5
52 identifiable discussions cover 69/75 comments; the largest has 5. Remaining thread identities are unknown; this is not a count of independent users.
Selected individual opinions
One user fed both models a script meant to print the max of two numbers; Opus 5 deduced it correctly, while Opus 5.5 wrongly concluded it prints 1/0. One user felt Opus 5 was degraded today, with no solid data, just noticing it made more mistakes than usual.
About the sample and scores
As of 2026-10-02 22:02:03 UTC, Claude Opus 5 in the Claude family has 75 explicitly attributed community comments in the last 7 days; its community score is 20.8/100. Available category scores include General text 19.0/100 (n=46); Reasoning 36.8/100 (n=16); Speed & latency 34.2/100 (n=6). Scoring method and sources
RECENT EXPERIENCE
How the Experience Index is changing
30 days of rolling 7-day scores at each daily snapshot. The vertical scale adapts to the data. Missing or insufficient samples leave gaps; today is still updating.
Tap or use ← → for dates, scores and sample sizes. Blank dates have insufficient data.
Daily readings & sample sizes
| Date | Score | n | Scoring window (UTC) |
|---|---|---|---|
| 2026-10-02 | 39.2 | 6 | 2026-09-25T22:02:03+00:00 – 2026-10-02T22:02:03+00:00 |
| 2026-10-01 | 39.2 | 6 | 2026-09-24T23:57:02+00:00 – 2026-10-01T23:57:02+00:00 |
| 2026-09-30 | 39.2 | 6 | 2026-09-23T23:57:03+00:00 – 2026-09-30T23:57:03+00:00 |
| 2026-09-29 | 34.4 | 8 | 2026-09-22T23:57:03+00:00 – 2026-09-29T23:57:03+00:00 |
| 2026-09-28 | 19.5 | 23 | 2026-09-21T23:57:03+00:00 – 2026-09-28T23:57:03+00:00 |
| 2026-09-27 | 20.4 | 26 | 2026-09-20T23:57:03+00:00 – 2026-09-27T23:57:03+00:00 |
| 2026-09-26 | 20.4 | 26 | 2026-09-19T23:57:03+00:00 – 2026-09-26T23:57:03+00:00 |
| 2026-09-25 | 21.1 | 25 | 2026-09-18T23:57:03+00:00 – 2026-09-25T23:57:03+00:00 |
| 2026-09-24 | 21.1 | 25 | 2026-09-17T23:57:03+00:00 – 2026-09-24T23:57:03+00:00 |
| 2026-09-23 | 21.9 | 24 | 2026-09-16T23:57:03+00:00 – 2026-09-23T23:57:03+00:00 |
| 2026-09-22 | 25.0 | 20 | 2026-09-15T23:57:03+00:00 – 2026-09-22T23:57:03+00:00 |
| 2026-09-21 | 38.7 | 6 | 2026-09-14T23:57:03+00:00 – 2026-09-21T23:57:03+00:00 |
| 2026-09-20 | — | 3 | 2026-09-13T23:57:03+00:00 – 2026-09-20T23:57:03+00:00 |
| 2026-09-19 | — | 3 | 2026-09-12T23:57:03+00:00 – 2026-09-19T23:57:03+00:00 |
| 2026-09-18 | 39.5 | 8 | 2026-09-11T23:57:02+00:00 – 2026-09-18T23:57:02+00:00 |
| 2026-09-17 | 37.8 | 10 | 2026-09-10T23:57:03+00:00 – 2026-09-17T23:57:03+00:00 |
| 2026-09-16 | 35.3 | 12 | 2026-09-09T23:57:03+00:00 – 2026-09-16T23:57:03+00:00 |
| 2026-09-15 | 35.4 | 12 | 2026-09-08T23:57:03+00:00 – 2026-09-15T23:57:03+00:00 |
| 2026-09-14 | 36.6 | 13 | 2026-09-07T23:57:03+00:00 – 2026-09-14T23:57:03+00:00 |
| 2026-09-13 | 33.7 | 12 | 2026-09-06T23:57:03+00:00 – 2026-09-13T23:57:03+00:00 |
| 2026-09-12 | 36.6 | 13 | 2026-09-05T23:57:03+00:00 – 2026-09-12T23:57:03+00:00 |
| 2026-09-11 | 40.5 | 9 | 2026-09-04T23:00:04+00:00 – 2026-09-11T23:00:04+00:00 |
| 2026-09-10 | 34.8 | 11 | 2026-09-03T23:00:05+00:00 – 2026-09-10T23:00:05+00:00 |
| 2026-09-09 | 44.7 | 16 | 2026-09-02T23:00:04+00:00 – 2026-09-09T23:00:04+00:00 |
| 2026-09-08 | 47.1 | 19 | 2026-09-01T23:00:03+00:00 – 2026-09-08T23:00:03+00:00 |
| 2026-09-07 | 41.3 | 23 | 2026-08-31T23:00:03+00:00 – 2026-09-07T23:00:03+00:00 |
| 2026-09-06 | 41.5 | 24 | 2026-08-30T23:00:02+00:00 – 2026-09-06T23:00:02+00:00 |
| 2026-09-05 | 39.5 | 23 | 2026-08-29T23:00:03+00:00 – 2026-09-05T23:00:03+00:00 |
| 2026-09-04 | 40.6 | 22 | 2026-08-28T23:00:03+00:00 – 2026-09-04T23:00:03+00:00 |
| 2026-09-03 | 42.3 | 22 | 2026-08-27T23:00:03+00:00 – 2026-09-03T23:00:03+00:00 |
Hacker News · 39.2 · 6 scored comments · Read platform reviews →
ONE MODEL, DIFFERENT COMMUNITIES
Across the communities
Scored separately for each platform, with different samples and audiences. Platform scores are not simply averaged; scored counts exclude quota-only comments.
Scores range from 0 to 100. Select a platform for its trend and reviews. Latest opinion time is not crawler health.
Explore long-term changes & analysis
Model lifecycle
Claude Opus 5
Released 2026-07-24 · Back to Claude
Exact-version signals observed 2026-09-21–2026-10-01 · n=121
Data window: 2026-09-04 00:00:00 to 2026-10-02 00:00:00 UTC · Scoring method: experience_score_v3.
Lifecycle updated daily · latest complete UTC day 2026-10-01
The current classification method has insufficient evidence in the original baseline window. The window is preserved; no long-term improvement or decline is inferred, and old-method scores are not compared.
- Fixed 14-day baseline
- – n=0
- Latest 28 days
- 21.9 n=114
- Change vs baseline
- –
- 90-day slope
- – pts / 30 days
14-day rolling experience
Each daily point summarizes the previous 14 complete UTC days. The dashed line is the fixed release baseline; the verdict still uses independent non-overlapping 14-day segments.
0 of 4 independent 14-day segments ready for trend judgment.
What explains the change
Experience change and discussion-mix change are separated. These are observational contributions, not proof of cause.
No category has enough samples in both windows for a contribution conclusion.
| Category | Baseline → current | Category change | Weight share | Experience contribution | Discussion-mix contribution |
|---|---|---|---|---|---|
| Coding | Baseline → currentnot enough datan=0 → 9 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Image generation | Baseline → currentnot enough datan=0 → 0 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Image understanding | Baseline → currentnot enough datan=0 → 0 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Local deploy | Baseline → currentnot enough datan=0 → 0 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Reasoning | Baseline → currentnot enough datan=0 → 26 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Roleplay / creative | Baseline → currentnot enough datan=0 → 0 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Safety & refusals | Baseline → currentnot enough datan=0 → 5 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Speed & latency | Baseline → currentnot enough datan=0 → 8 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| General text | Baseline → currentnot enough datan=0 → 66 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
| Video generation | Baseline → currentnot enough datan=0 → 0 | Category change– | Weight share– | Experience contribution– | Discussion-mix contribution– |
Lifecycle scores use the same experience-signal weights as the main index, without sample shrinkage after n=30. They measure public user perception, not model capability or backend causes.
Community reviews
Selected comments from the last 7 days. The balance of excerpts does not represent the share of positive reviews.
1 selected excerpts
Community reports about limits cannot establish official allowances.
Comment context
For example if you have Opus 5 process a bunch of reviews from an subagent fanout that is reviewing code, then synthesizing those into a final report. It keeps more of the actual defects vs 5.5, by a small margin. But it also keeps more false positive and takes about twice as long and twice as many tokens as 5.5
Comment context
But right around the time my monthly limit reset this evening, I noticed an extreme shift in communication style and coding behavior that feels suspiciously like Opus 5 (O5). For the last week, O5.5 would take my requirements and immediately get to work, usually knocking out a feature in minutes and doing an incredible job across implementation, UX, architecture, and token efficiency. Tonight, I’m seeing something very different. The first tell was the communication style. O5.5 suddenly started responding to feature requests with these glazing pre-implementation narratives and generic estimates like, “That’s a great idea. That’s two days’ worth of work and it will be amazing,” and then proceeding to implement the feature in the most slopcode-slathered way possible. That was a very specific and particularly annoying characteristic of O5, so I clocked it immediately. The next tell: a sudden reversion to walls-of-text responses when prodded for details on...*anything*.
Comment context
I feed it a script which takes two numbers and prints the max of the two numbers. Opus 5.5 thinks it prints 1/0 instead of the max numbers.
Show original
opus5太慢了
Comment context
Actually I thought Opus 5 was great, and don't feel like 5.5 is some incredible leap forward from it
Comment context
Tried to use muse for PR review and opus 5 was hating on the feedback and instead opus found a different real issue that muse didn't.
No selected comments for this filter in the last 7 days.