← All models

GPT · Last 7 days

GPT-6 Luna

0–100 · higher means more positive community experience, not benchmark performance.

Updated 2026-10-01 09:32 UTC · 2026-09-24 – 2026-10-01 UTC
Experience score35.5/100114 comments in 7 days

What people think

One user notes that on the DeepSWE v1.1 software engineering test, GPT-6 Luna scored 66.6%, near Opus 5/Fable 5 medium effort, with per-task cost ~93%/96% lower. One user asked GPT-6 Luna to use computer use to solve a problem, but it stubbornly refused despite repeated reminders

About the sample and scores

As of 2026-10-01 09:52:04 UTC, GPT-6 Luna in the GPT family has 114 explicitly attributed community comments in the last 7 days; its community score is 35.5/100. Available category scores include General text 36.6/100 (n=52); Coding 34.9/100 (n=14); Reasoning 40.4/100 (n=16). Scoring method and sources

Community reviews

Selected comments from the last 7 days. The balance of excerpts does not represent the share of positive reviews.

15 selected excerpts

Community reports about limits cannot establish official allowances.

PositiveCoding

“GPT-6 Luna scored 66.6%, close to Opus 5 and Fable 5's medium effort, while the per-task cost is about 93% and 96% lower”

Machine translated
Show original

GPT-6 Luna 得分为 66.6%,接近 Opus 5 和 Fable 5 的 medium effort,单任务成本则低约 93% 和 96%

Comment context
From the same comment ↗

Software engineering benchmarks tell a clearer story. In DeepSWE v1.1, GPT-6 Sol's max effort score is 68.8%, just 1.1 percentage points behind Claude Fable 5's top score of 69.9%, but with a per-task cost about 80% lower. GPT-6 Luna scores 66.6%, close to Opus 5 and Fable 5 at medium effort, with per-task costs about 93% and 96% lower respectively.

Machine translated
Show original

软件工程测试更能说明问题。在 DeepSWE v1.1 中,GPT-6 Sol 的 max effort 得分为 68.8%,距离 Claude Fable 5 的最高成绩 69.9% 只差 1.1 个百分点,但单任务成本低约 80%。GPT-6 Luna 得分为 66.6%,接近 Opus 5 和 Fable 5 的 medium effort,单任务成本则低约 93% 和 96%。

NegativeSafety & refusals

“Asked it to use computeruse to solve my problem, reminded it for ages but it just wouldn't dare”

Machine translated
Show original

叫它用computeruse帮我解决问题,提醒半天死活不敢

PositiveUsage limits

“in pro x5 its been used for goal 1 week straight heavily and only used 3-5%”

NegativeSpeed & latency

“6luna is both stupid and slow”

Machine translated
Show original

6luna是个又蠢又慢的模型

NegativeCoding

“just a "done!" where you then have to clean everything up afterwards”

Comment context
From the same comment ↗

I've had bad experiences with gpt 6 luna, but 5.6 luna worked real nicely with e.g React, you can do a back n forth and it'll work like a teammate vs just a "done!" where you then have to clean everything up afterwards

NegativeSpeed & latency

“You have to set it to the highest setting and go back and forth two or three rounds before it produces anything.”

Machine translated
Show original

要开最高档,反复沟通两三轮才能出一点东西

NeutralSpeed & latency

“The speed seems to be a bit faster than it was in 5.6”

Machine translated
Show original

速度貌似是比5.6时候快了一些

PositiveReasoning

“Clearly stronger in capability than DeepSeekv4f and pro”

Machine translated
Show original

能力上明确强于DeepSeekv4f和pro

NegativeReasoning

“It's slightly worse, but I think as a sub-agent it doesn't matter—just modify the main agent's prompt to have it break down complex tasks into finer pieces.”

Machine translated
Show original

稍差一点,但是我觉得当子智能体的话不耽误啊,改一下主智能体的提示词,让它把复杂任务拆细就是了。

Comment context
Comment being replied to ↗

I haven't used it much recently. Is 6 luna a regression compared to 5.6 luna? I previously set the sub-agent to force 6 luna max; do I need to change it back to 5.6 luna max? [crying laughing]

Machine translated
Show original

我最近没怎么用,6 luna跟5.6luna比难道是退步的么?我之前是设置了子代理强制用6luna max,需要改回5.6 luna max么[笑哭]

PositiveReasoning

“API answers are way better because I use Luna kind of like an LLM director in a game”

NegativeRoleplay / creative

“I don't see any difference”

Machine translated
Show original

我看没什么区别啊

Comment context
From the same comment ↗

Is there any substantive improvement in 6luna compared to 5.6luna? I don't see any difference.

Machine translated
Show original

6luna和5.6luna有什么实质性提升吗?我看没什么区别啊

PositiveCoding

“Specially recommend GPT6 luna, generous in quantity, performance not much different from sol”

Machine translated
Show original

特别推荐GPT6 luna,量大管饱,性能跟sol差距不大

NegativeUsage limits

“Now with Luna 6 I measured after full spending of my weekly, 44 dolars. This 50% price reduction is sounding like a total scam.”

How's your AI experience today?