← 返回全部分类

当前排名

推理

截至 2026-09-15 06:57:03 UTC,近 7 日推理体感排名:闭源组 — Claude Fable 5(86.5 分,n=66);开源组 — GLM 5.3 Flash(79.3 分,n=38)。

每个场景展示 Top 3 具体版本,样本量满 5 条即可上榜;排名变化对比上一次发布的榜单。
#1 Claude Fable 5n=66 86.5

“Fable基本one shot everything”

来自 知乎 · 2026-09-15

“These things you mentioned can be somehow avoided using less reasoning and more workflow (skills and etc) but that comes with a cost”

来自 Reddit · 英文 · 2026-09-14
查看模型详情 →
#2 GPT-6 Astran=156 79.7

“Astra is basically S+ tier and Fable is S tier; it can do stuff that Fable flubs but you really need a project that requires that much god-tier brainpower.”

来自 Reddit · 英文 · 2026-09-15

“pretty much not up for debate that Astra is the best model in the world”

来自 Reddit · 英文 · 2026-09-15
查看模型详情 →
#3 Claude Fable 5.1n=32 77.1

“推理不如Fable5.1”

来自 知乎 · 2026-09-15

“you end up doing less rework since Fable agents are just much smarter”

来自 Hacker News · 英文 · 2026-09-14
查看模型详情 →
#1 GLM 5.3 Flashn=38 79.3

“it can't even consistently complete my annotation jobs without making up its own output schema or wrapping the json with code fences”

来自 Reddit · 英文 · 2026-09-15

“GLM 5.3 flash on the other hand is... Special...”

来自 Reddit · 英文 · 2026-09-14
查看模型详情 →
#2 Qwen3.8 27Bn=43 79.1

“qwen3.8 27b is pretty mind blowing to me. It can easily solve portswigger labs for example at q4_k_m.”

来自 Hacker News · 英文 · 2026-09-15

“Qwen3.8 27B for reasoning, coding, and agentic work”

来自 Reddit · 英文 · 2026-09-15
查看模型详情 →
#3 DeepSeek V4 Flashn=17 74.1

“DSV4's implementation was broken beyond 90k context”

来自 Reddit · 英文 · 2026-09-14

“DeepSeek Flash takes the spotlight”

来自 Reddit · 英文 · 2026-09-14
查看模型详情 →