← All models

GPT · Last 7 days

GPT-6 Luna

0–100 · higher means more positive community experience, not benchmark performance.

Updated 2026-10-01 09:57 UTC · 2026-09-24 – 2026-10-01 UTC
Experience score34.6/100113 comments in 7 days

What people think

One user notes that on the DeepSWE v1.1 software engineering test, GPT-6 Luna scored 66.6%, near Opus 5/Fable 5 medium effort, with per-task cost ~93%/96% lower. One user asked GPT-6 Luna to use computer use to solve a problem, but it stubbornly refused despite repeated reminders

About the sample and scores

As of 2026-10-01 10:42:02 UTC, GPT-6 Luna in the GPT family has 113 explicitly attributed community comments in the last 7 days; its community score is 34.6/100. Available category scores include General text 36.6/100 (n=52); Coding 34.9/100 (n=14); Reasoning 36.6/100 (n=15). Scoring method and sources

Community reviews

Selected comments from the last 7 days. The balance of excerpts does not represent the share of positive reviews.

6 selected excerpts

Community reports about limits cannot establish official allowances.

PositiveCoding

“GPT-6 Luna scored 66.6%, close to Opus 5 and Fable 5's medium effort, while the per-task cost is about 93% and 96% lower”

Machine translated
Show original

GPT-6 Luna 得分为 66.6%,接近 Opus 5 和 Fable 5 的 medium effort,单任务成本则低约 93% 和 96%

Comment context
From the same comment ↗

Software engineering benchmarks tell a clearer story. In DeepSWE v1.1, GPT-6 Sol's max effort score is 68.8%, just 1.1 percentage points behind Claude Fable 5's top score of 69.9%, but with a per-task cost about 80% lower. GPT-6 Luna scores 66.6%, close to Opus 5 and Fable 5 at medium effort, with per-task costs about 93% and 96% lower respectively.

Machine translated
Show original

软件工程测试更能说明问题。在 DeepSWE v1.1 中,GPT-6 Sol 的 max effort 得分为 68.8%,距离 Claude Fable 5 的最高成绩 69.9% 只差 1.1 个百分点,但单任务成本低约 80%。GPT-6 Luna 得分为 66.6%,接近 Opus 5 和 Fable 5 的 medium effort,单任务成本则低约 93% 和 96%。

PositiveUsage limits

“in pro x5 its been used for goal 1 week straight heavily and only used 3-5%”

PositiveReasoning

“Clearly stronger in capability than DeepSeekv4f and pro”

Machine translated
Show original

能力上明确强于DeepSeekv4f和pro

PositiveReasoning

“API answers are way better because I use Luna kind of like an LLM director in a game”

PositiveCoding

“Specially recommend GPT6 luna, generous in quantity, performance not much different from sol”

Machine translated
Show original

特别推荐GPT6 luna,量大管饱,性能跟sol差距不大

How's your AI experience today?