← All categories

Current ranking

Reasoning

For the 7-day window ending 2026-09-13 23:17:02 UTC, user-experience leaders in Reasoning: closed models — Claude Fable 5 (82.3/100, n=64); open models — Qwen3.8 27B (80.0/100, n=45).

Top 3 specific models in each category. A model appears when its category sample size is at least 5. Rank movement compares the previous published snapshot.
#1 Claude Fable 5n=64 82.3

“Astra is very, very good. So is Fable. I think we're still getting a feel for what each is better at. There are definitely use cases where Astra beats Fable and vice versa”

via Reddit · 2026-09-13

“Fable ends up saying Astra is right and it missed x because it had either made some assumptions or not gone deep enough”

via Reddit · 2026-09-13
Open model detail →
#2 GPT-6 Astran=189 81.0

“Astra is no match for very architecturally or algorithmically complex task”

via Reddit · 2026-09-13
Open model detail →
#3 Claude Fable 5.1n=25 74.4

“I think Fable 5.1 was nerfed and I cannot trust it anymore on this kind of logic”

via Reddit · 2026-09-13

“Fable seems to forget about things”

via Reddit · 2026-09-10
Open model detail →
#1 Qwen3.8 27Bn=45 80.0

“doesn't overthinking like 3.8 does on xhigh”

via Reddit · 2026-09-13

“Qwen3.8:27b doesnt know when to not think. It overthinks things that are simple and thinks hard for big tasks too (appropriately).”

via Reddit · 2026-09-13
Open model detail →
#2 GLM 5.3 Flashn=34 79.8

“The flash version is very capable and fits into the luna category. Feels a little more capable than luna imo but more verbose.”

via Reddit · 2026-09-13

“I think the best here is glm 5.3 flash, it has a very low hallucination rate”

via Reddit · 2026-09-13
Open model detail →
#3 DeepSeek V4 Flashn=18 75.7

“DSV4 it will just keep on going until it finds the issue, or if it cannot, then it'll stop eventually and prepare a handover for a fresh session”

via Reddit · 2026-09-13

“never have issues with agents going rogue like this”

via Reddit · 2026-09-13
Open model detail →