GPT · Last 7 days

GPT-5.6 Luna

0–100 · higher means more positive community experience, not benchmark performance.

Updated 2026-10-04 07:42 UTC · 2026-09-27 – 2026-10-04 UTC
All-platform score61.0/10031 comments in 7 days

Discussion overview

What the full sample discusses

All platforms · 31 eligible comments over 7 days, including usage-limit feedback; deduplicated by source and comment.

  • General text · 11 comments: 7 positive / 4 negative / 0 mixed or neutral
  • Usage limits · 8 comments: 5 positive / 2 negative / 1 mixed or neutral
  • Reasoning · 6 comments: 4 positive / 2 negative / 0 mixed or neutral
  • Coding · 4 comments: 3 positive / 1 negative / 0 mixed or neutral

Comments may cover multiple dimensions; counts are not additive or equivalent to the weighted score. Links show selected source excerpts.

Recent change

Rolling 7-day score: +13.7 points vs 7 days ago. Discussion mix can affect scores; this does not establish a cause.

Source coverage and discussion concentration

Hacker News 3 · Reddit 21 · Zhihu 7

21 identifiable discussions cover 28/31 comments; the largest has 3. Remaining thread identities are unknown; this is not a count of independent users.

Selected individual opinions

One user feels that the quota of ChatGPT's free model 5.6 Luna is relatively sufficient One user reported that GPT-5.6 Luna was extremely unintelligent, unable to use tool calls, and shockingly poor in chat.

About the sample and scores

As of 2026-10-04 07:42:03 UTC, GPT-5.6 Luna in the GPT family has 31 explicitly attributed community comments in the last 7 days; its community score is 61.0/100. Available category scores include General text 58.1/100 (n=11); Reasoning 57.8/100 (n=6); Usage limits 59.1/100 (n=8). Scoring method and sources

RECENT EXPERIENCE

How the Experience Index is changing

Not enough historyvs 7 days ago · points

30 days of rolling 7-day scores at each daily snapshot. The vertical scale adapts to the data. Missing or insufficient samples leave gaps; today is still updating.

Tap or use ← → for dates, scores and sample sizes. Blank dates have insufficient data.

No trend available

RedNote · Not enough opinions · 0 scored comments · Read platform reviews →

ONE MODEL, DIFFERENT COMMUNITIES

Across the communities

Last 7 days

Scored separately for each platform, with different samples and audiences. Platform scores are not simply averaged; scored counts exclude quota-only comments.

Scores range from 0 to 100. Select a platform for its trend and reviews. Latest opinion time is not crawler health.

Community reviews

Selected comments from the last 7 days. The balance of excerpts does not represent the share of positive reviews.

0 selected excerpts

Community reports about limits cannot establish official allowances.

No selected comments for this filter in the last 7 days.

Explore long-term changes & analysis

Model lifecycle

GPT-5.6 Luna

Released 2026-07-10 · Back to GPT

Exact-version signals observed 2026-09-23–2026-10-03 · n=74

Data window: 2026-09-06 00:00:00 to 2026-10-04 00:00:00 UTC · Scoring method: experience_score_v3.

Insufficient historical baseline The original baseline window lacks enough current-method evidence for a long-term comparison.

Lifecycle updated daily · latest complete UTC day 2026-10-03

The current classification method has insufficient evidence in the original baseline window. The window is preserved; no long-term improvement or decline is inferred, and old-method scores are not compared.

Fixed 14-day baseline
–
n=0
Latest 28 days
58.5
n=60
Change vs baseline
–
90-day slope
–
pts / 30 days

14-day rolling experience

Each daily point summarizes the previous 14 complete UTC days. The dashed line is the fixed release baseline; the verdict still uses independent non-overlapping 14-day segments.

14-day rolling experience4050602026-09-16–2026-09-30 · 54.8 · n=492026-09-17–2026-10-01 · 57.5 · n=532026-09-18–2026-10-02 · 57.7 · n=562026-09-19–2026-10-03 · 58.4 · n=582026-09-20–2026-10-04 · 58.5 · n=6008-1308-2208-3109-0909-1809-2710-04

0 of 4 independent 14-day segments ready for trend judgment.

What explains the change

Experience change and discussion-mix change are separated. These are observational contributions, not proof of cause.

No category has enough samples in both windows for a contribution conclusion.

Category Baseline → current Category change Weight share Experience contribution Discussion-mix contribution
Coding Baseline → currentnot enough datan=0 → 9 Category change– Weight share– Experience contribution– Discussion-mix contribution–
Image generation Baseline → currentnot enough datan=0 → 0 Category change– Weight share– Experience contribution– Discussion-mix contribution–
Image understanding Baseline → currentnot enough datan=0 → 0 Category change– Weight share– Experience contribution– Discussion-mix contribution–
Local deploy Baseline → currentnot enough datan=0 → 0 Category change– Weight share– Experience contribution– Discussion-mix contribution–
Reasoning Baseline → currentnot enough datan=0 → 12 Category change– Weight share– Experience contribution– Discussion-mix contribution–
Roleplay / creative Baseline → currentnot enough datan=0 → 0 Category change– Weight share– Experience contribution– Discussion-mix contribution–
Safety & refusals Baseline → currentnot enough datan=0 → 0 Category change– Weight share– Experience contribution– Discussion-mix contribution–
Speed & latency Baseline → currentnot enough datan=0 → 6 Category change– Weight share– Experience contribution– Discussion-mix contribution–
General text Baseline → currentnot enough datan=0 → 34 Category change– Weight share– Experience contribution– Discussion-mix contribution–
Video generation Baseline → currentnot enough datan=0 → 0 Category change– Weight share– Experience contribution– Discussion-mix contribution–

Lifecycle scores use the same experience-signal weights as the main index, without sample shrinkage after n=30. They measure public user perception, not model capability or backend causes.

How's your AI experience today?