Claude · Last 7 days

Claude Opus 5

0–100 · higher means more positive community experience, not benchmark performance.

Updated 2026-10-04 09:22 UTC · 2026-09-27 – 2026-10-04 UTC
All-platform score23.9/10064 comments in 7 days

Discussion overview

What the full sample discusses

All platforms · 64 eligible comments over 7 days, including usage-limit feedback; deduplicated by source and comment.

  • General text · 38 comments: 4 positive / 29 negative / 5 mixed or neutral
  • Reasoning · 15 comments: 5 positive / 10 negative / 0 mixed or neutral
  • Usage limits · 5 comments: 2 positive / 2 negative / 1 mixed or neutral
  • Speed & latency · 4 comments: 0 positive / 3 negative / 1 mixed or neutral

Comments may cover multiple dimensions; counts are not additive or equivalent to the weighted score. Links show selected source excerpts.

Recent change

Rolling 7-day score: −3.2 points vs 7 days ago. Discussion mix can affect scores; this does not establish a cause.

Source coverage and discussion concentration

Hacker News 7 · Reddit 50 · v2ex 1 · RedNote 1 · Zhihu 5

47 identifiable discussions cover 57/64 comments; the largest has 5. Remaining thread identities are unknown; this is not a count of independent users.

Selected individual opinions

One user fed both models a script meant to print the max of two numbers; Opus 5 deduced it correctly, while Opus 5.5 wrongly concluded it prints 1/0. One user replied that even after the nerf, Opus 5.5 is still much better than Opus 5.

About the sample and scores

As of 2026-10-04 09:22:04 UTC, Claude Opus 5 in the Claude family has 64 explicitly attributed community comments in the last 7 days; its community score is 23.9/100. Available category scores include General text 22.1/100 (n=38); Reasoning 38.3/100 (n=15); Usage limits 49.3/100 (n=5). Scoring method and sources

RECENT EXPERIENCE

How the Experience Index is changing

Not enough historyvs 7 days ago · points

30 days of rolling 7-day scores at each daily snapshot. The vertical scale adapts to the data. Missing or insufficient samples leave gaps; today is still updating.

Tap or use ← → for dates, scores and sample sizes. Blank dates have insufficient data.

No trend available

RedNote · Not enough opinions · 1 scored comments · Read platform reviews →

ONE MODEL, DIFFERENT COMMUNITIES

Across the communities

Last 7 days

Scored separately for each platform, with different samples and audiences. Platform scores are not simply averaged; scored counts exclude quota-only comments.

Scores range from 0 to 100. Select a platform for its trend and reviews. Latest opinion time is not crawler health.

Community reviews

Selected comments from the last 7 days. The balance of excerpts does not represent the share of positive reviews.

0 selected excerpts

Community reports about limits cannot establish official allowances.

No selected comments for this filter in the last 7 days.

Explore long-term changes & analysis

Model lifecycle

Claude Opus 5

Released 2026-07-24 · Back to Claude

Exact-version signals observed 2026-09-21–2026-10-03 · n=131

Data window: 2026-09-06 00:00:00 to 2026-10-04 00:00:00 UTC · Scoring method: experience_score_v3.

Insufficient historical baseline The original baseline window lacks enough current-method evidence for a long-term comparison.

Lifecycle updated daily · latest complete UTC day 2026-10-03

The current classification method has insufficient evidence in the original baseline window. The window is preserved; no long-term improvement or decline is inferred, and old-method scores are not compared.

Fixed 14-day baseline
–
n=0
Latest 28 days
21.6
n=123
Change vs baseline
–
90-day slope
–
pts / 30 days

14-day rolling experience

Each daily point summarizes the previous 14 complete UTC days. The dashed line is the fixed release baseline; the verdict still uses independent non-overlapping 14-day segments.

14-day rolling experience20304050602026-09-13–2026-09-27 · 25.2 · n=682026-09-14–2026-09-28 · 23.5 · n=772026-09-15–2026-09-29 · 22.7 · n=932026-09-16–2026-09-30 · 23.0 · n=1052026-09-17–2026-10-01 · 21.6 · n=1112026-09-18–2026-10-02 · 22.3 · n=1182026-09-19–2026-10-03 · 21.6 · n=12108-1308-2208-3109-0909-1809-2710-03

0 of 4 independent 14-day segments ready for trend judgment.

What explains the change

Experience change and discussion-mix change are separated. These are observational contributions, not proof of cause.

No category has enough samples in both windows for a contribution conclusion.

Category Baseline → current Category change Weight share Experience contribution Discussion-mix contribution
Coding Baseline → currentnot enough datan=0 → 10 Category change– Weight share– Experience contribution– Discussion-mix contribution–
Image generation Baseline → currentnot enough datan=0 → 0 Category change– Weight share– Experience contribution– Discussion-mix contribution–
Image understanding Baseline → currentnot enough datan=0 → 0 Category change– Weight share– Experience contribution– Discussion-mix contribution–
Local deploy Baseline → currentnot enough datan=0 → 0 Category change– Weight share– Experience contribution– Discussion-mix contribution–
Reasoning Baseline → currentnot enough datan=0 → 29 Category change– Weight share– Experience contribution– Discussion-mix contribution–
Roleplay / creative Baseline → currentnot enough datan=0 → 0 Category change– Weight share– Experience contribution– Discussion-mix contribution–
Safety & refusals Baseline → currentnot enough datan=0 → 6 Category change– Weight share– Experience contribution– Discussion-mix contribution–
Speed & latency Baseline → currentnot enough datan=0 → 9 Category change– Weight share– Experience contribution– Discussion-mix contribution–
General text Baseline → currentnot enough datan=0 → 70 Category change– Weight share– Experience contribution– Discussion-mix contribution–
Video generation Baseline → currentnot enough datan=0 → 0 Category change– Weight share– Experience contribution– Discussion-mix contribution–

Lifecycle scores use the same experience-signal weights as the main index, without sample shrinkage after n=30. They measure public user perception, not model capability or backend causes.

How's your AI experience today?