当前最佳模型
按任务场景看现在哪个模型好用——基于最近 7 天真实用户体验排名。
与模型详情页同口径 · 7 日窗口 · 更新于 2026-09-11 17:00 UTC每个场景展示 Top 3 具体版本,样本量满 5 条即可上榜;排名变化对比上一次发布的榜单。 有匹配评论时可展开查看;暂无评论的版本直接跳转模型详情页。
通用文本
#1 Claude Opus 4.6n=8 72.7▾
#2 Gemini 3.8 Flashn=17 62.4▾
#3 GPT-5.5n=9 57.8▾
“GPT-5.5 at the same reasoning level does not have this. Same account, same standard chat, same prompt, minutes apart: 5.6 Sol fails every time, 5.5 completes and actually runs the search.”
来自 Reddit · 英文 · 2026-09-09
查看模型详情 →“A good plan makes 5.5 sufficient; a better model may reduce how much planning and supervision you need.”
来自 Reddit · 英文 · 2026-09-08
编程
#2 Claude Fable 5n=74 67.5▾
推理
#1 Claude Fable 5n=72 81.1▾
“Fable is very very good at long context through repeated compactions. With a proper handoff file it follows along very very well.”
来自 Reddit · 英文 · 2026-09-11
查看模型详情 →“Alpöge himself presented a counterexample to the Jacobian conjecture on July, found with Claude Fable”
来自 Hacker News · 英文 · 2026-09-11
#2 GPT-6 Astran=266 81.0▾
“Same thing happened to me. When I asked it to stop, it started engineering an elaborate rollback process, talking about what would happen if two people tried rolling back simultaneously, etc.”
来自 Reddit · 英文 · 2026-09-11
查看模型详情 →“Astra is way more capable and seems more like a genuine collaborator that Opus pretends to be and yeah it also gets things done a bit better”
来自 Reddit · 英文 · 2026-09-11
图像/视觉
#1 GPT-5.4 Image 2n=13 66.6▾
“generates more realistic images than gemini”
来自 Hacker News · 英文 · 2026-09-09
查看模型详情 →“明显是2的意境更好”
来自 知乎 · 2026-09-09
#2 GPT Image 2.5n=50 66.3▾
角色扮演/创作
#1 GPT-6 Astran=14 47.0▾
#2 GPT-5.6 Soln=5 38.9▾
#3 Grok 4.6n=5 36.1▾
“Newer ones have every character speak in complete, helpful, on-topic sentences — because that's what the assistant training rewards. Everyone sounds like the same person wearing different hats.”
来自 Reddit · 英文 · 2026-09-08
查看模型详情 →“4 and 4.1 were amazing. 4.6 is useless. Purely a downgrade.”
来自 Reddit · 英文 · 2026-09-08
速度与延迟
#1 GPT Image 2.5n=7 72.4▾
#2 GPT-5.6 Lunan=6 68.2▾
#3 Gemini 3.8 Flashn=13 63.0▾
每个场景展示 Top 3 具体版本,样本量满 5 条即可上榜;排名变化对比上一次发布的榜单。 有匹配评论时可展开查看;暂无评论的版本直接跳转模型详情页。
通用文本
#1 GLM 5.2n=17 75.8▾
#2 GLM 5.3 Flashn=59 74.5▾
#3 DeepSeek V4 Flashn=72 62.6▾
“pro v4 was great, this switch to flash has stripped a lot of nuance from it.. It's become very blunt, direct and even aggressive. The analyis and ideas are actually strong though, better than before, very good on that front, but socially it”
来自 Reddit · 英文 · 2026-09-10
查看模型详情 →“直接连v4好像连不上了”
来自 知乎 · 2026-09-10
编程
#1 GLM 5.3 Flashn=23 74.3▾
#2↑1 DeepSeek V4 Flashn=13 67.1▾
“They handle most everyday coding-agent tasks fine, and you can run them usage-based or on a flat plan”
来自 Reddit · 英文 · 2026-09-11
查看模型详情 →“it should have no problem adapting to your codebase. Get yourself an API key, install "dsh", put in the key, and get coding with deepseek-v4-flash”
来自 Reddit · 英文 · 2026-09-10
推理
#1 Qwen3.8 27Bn=52 79.3▾
“going from 3.6-27B to 3.8.27B is that it works much better to solve problems over time, figuring out how to gather data, trying multiple approaches, not getting down an endless rabbit hole, etc.”
来自 Reddit · 英文 · 2026-09-11
查看模型详情 →“I'm Qwen3.8 27B to some times overthink itself to the wrong answer, will have to compare”
来自 Reddit · 英文 · 2026-09-10
#2 GLM 5.3 Flashn=31 78.5▾
速度与延迟
#1 DeepSeek V4 Flashn=66 71.3▾
#2 Qwen3.6 35B A3Bn=12 67.7▾
#3 Qwen3.8 Flash Nextn=57 67.6▾
“I even managed to get IQ\_4\_XS Qwen3.8 Flash Next running at \~15tok/s and 65k context which is wild for a 100GB model file (admittedly it ate up every single MB of my 64GB of RAM and even my SSD was at a solid 50% utilization at points so”
来自 Reddit · 英文 · 2026-09-11
查看模型详情 →“25-30 tokens per second with this config - 5090 & 256GB DDR5”
来自 Reddit · 英文 · 2026-09-11