Claude · Last 7 days
Claude Opus 4
0–100 · higher means more positive community experience, not benchmark performance.
Updated 2026-10-03 13:42 UTC · 2026-09-26 – 2026-10-03 UTCDiscussion overview
What the full sample discusses
All platforms · 1 eligible comments over 7 days, including usage-limit feedback; deduplicated by source and comment.
- Reasoning · 1 comments: 0 positive / 1 negative / 0 mixed or neutral
Comments may cover multiple dimensions; counts are not additive or equivalent to the weighted score. Links show selected source excerpts.
Recent change
Recent change: not enough history to assess improvement or decline.
Source coverage and discussion concentration
Zhihu 1
1 identifiable discussions cover 1/1 comments; the largest has 1. Remaining thread identities are unknown; this is not a count of independent users.
Selected individual opinions
One user notes that in the introspection test, even the best Claude Opus 4 and 4.1 only correctly detect and identify an internally injected thought about 20% of the time.
About the sample and scores
As of 2026-10-03 13:42:04 UTC, Claude Opus 4 in the Claude family has 1 explicitly attributed community comments in the last 7 days; its overall score is awaiting sufficient feedback. Scoring method and sources
RECENT EXPERIENCE
How the Experience Index is changing
30 days of rolling 7-day scores at each daily snapshot. The vertical scale adapts to the data. Missing or insufficient samples leave gaps; today is still updating.
Tap or use ← → for dates, scores and sample sizes. Blank dates have insufficient data.
No trend available
Reddit · Not enough opinions · 0 scored comments · Read platform reviews →
ONE MODEL, DIFFERENT COMMUNITIES
Across the communities
Scored separately for each platform, with different samples and audiences. Platform scores are not simply averaged; scored counts exclude quota-only comments.
Scores range from 0 to 100. Select a platform for its trend and reviews. Latest opinion time is not crawler health.
Explore long-term changes & analysis
This version is not yet tracked since release; recent 7-day experience is shown above.
Community reviews
Selected comments from the last 7 days. The balance of excerpts does not represent the share of positive reviews.
0 selected excerpts
Show original
表现最好的 Claude Opus 4 和 4.1,也只有大约 20% 的时候能察觉并说对
Comment context
Introspection, in plain terms, is whether you can see what's going on in your own mind. This can't be tested through conversation alone, so the researchers used a clever method. They first recorded the difference in a model's internal state when reading an all-caps text versus a normal text, essentially extracting a "shouting" thought; then they directly inserted this thought into the model's running internal state and asked it: Do you notice a thought being inserted? If it says "seems related to shouting" before this thought affects the output, it can only be because it looked at its own internals, not because it deduced it from the spoken words.
Machine translatedShow original
内省,说白了就是能不能看见自己脑子里在想什么。这事光靠聊天测不出来,所以论文用了个很巧的办法。研究员先录下模型读一段全大写文字和读一段普通文字时内部状态的差别,相当于抽出一个「大喊大叫」的念头;再把这个念头直接塞进模型正在运行的内部,然后问它:有没有察觉到被塞进来一个念头?如果它在这个念头还没影响到输出之前,就说出「好像跟大声喊叫有关」,那只能是看了自己的内部,不是从说出口的字里倒推出来的。
No selected comments for this filter in the last 7 days.