Methodology

“public-community experience index, not a benchmark”

What this index measures

The experience score is an observational index of how users in public communities describe their experience with each AI model family — positive vs negative experience signals extracted from real comments. It is not a capability benchmark and says nothing about what a model can or cannot do.

Data sources

We analyze a targeted sample of public, AI-related discussions from Reddit, Hacker News, Zhihu, RedNote (Xiaohongshu), and V2EX. This is not a census of any platform or a representative population survey. Each comment is classified by an LLM pipeline into experience signals per model and per dimension (coding, reasoning, safety refusals, speed…). Every valid opinion has equal weight in the global score; platform labels are retained for breakdowns and audit, not platform weighting.

Reddit

Platform. A network of communities organized around shared interests, where people post, vote, and comment.

Typical voices. People in AI, model, coding, product, and adjacent interest communities, often sharing hands-on comparisons of tools and models.

Why included. It adds broad English-language, practical, cross-model experience from many distinct communities.

Hacker News

Platform. A technology and startup discussion community centered on substantive conversation and intellectual curiosity.

Typical voices. Engineers, founders, researchers, and other technically engaged readers discussing systems and products.

Why included. It adds implementation detail and discussion of model behavior, release changes, reliability, latency, and cost.

Zhihu

Platform. A Chinese Q&A and content community built around sharing knowledge, experience, and insight.

Typical voices. Knowledge-seeking users, practitioners, and subject-matter contributors writing questions, answers, and longer explanations.

Why included. It adds explanatory Chinese-language accounts of output quality, pricing, workflows, and model-selection decisions.

RedNote

Platform. A life-interest community where people share lifestyle, learning, work, and creative experiences.

Typical voices. Everyday users, creators, students, and professionals describing AI in real tasks and routines.

Why included. RedNote adds adoption, usability, work and study workflows, creative use, and emotional reactions that technical forums may underrepresent. Only comments created after activation are eligible; older comments are not backfilled.

V2EX

Platform. A community for developers, designers, startup builders, and other internet-focused creators.

Typical voices. Technical practitioners discussing products, infrastructure, toolchains, and day-to-day operations.

Why included. It adds operational feedback on quality, latency, price, quotas, and integrations. Only experiences with official hosted model services are scoreable; third-party tools and local derivatives remain audit-only.

Windows and thresholds

The main score uses a 7-day rolling window; the pulse uses 24 hours. Model families or dimensions below the minimum sample threshold are shown as “not enough data” instead of a score. Every number on this site carries its sample size.

Version-level scores

A version score uses only comments that explicitly name that version. Generic mentions such as “ChatGPT” or “Claude” stay in the family score, so family sample sizes are larger. Versions below the minimum threshold show mention counts without a score.

Version lifecycle monitoring

Lifecycle pages use the first sample-sufficient 14-day window after release as a fixed baseline and compare it with the latest 28 complete UTC days. A daily 14-day rolling curve shows the observed path, while the formal trend verdict uses independent non-overlapping 14-day segments. Only exact, explicit version attribution is included. Category contributions separate changes within a category from changes in discussion mix; they explain the observed score movement but do not prove its cause.

Your direct reports

The one-tap report widget feeds a separate first-party signal. We deduplicate by hashed IP per model per day, never store raw IP addresses, and only display aggregate counts. Direct reports are cross-checked against community signals to detect manipulation; they do not enter the experience score itself.

Honesty rules

Changes are reported as user-perceived experience, never as claims about model capability. Anomaly annotations distinguish strong, weak, or no match with public events. When we don't know, we say so.