← 返回全部分类

当前排名

通用文本

截至 2026-09-13 21:17:03 UTC,近 7 日通用文本体感排名:闭源组 — Claude Opus 4.6(72.7 分,n=8);开源组 — GLM 5.2(73.2 分,n=14)。

每个场景展示 Top 3 具体版本,样本量满 5 条即可上榜;排名变化对比上一次发布的榜单。
#1 Claude Opus 4.6n=8 72.7

“Opus 4.6 was the last good model that could hold a conversation without devolving in to being an asshole while still being intelligent.”

来自 Reddit · 英文 · 2026-09-12

“massive love-fest for Opus 4.6. Y'all think it's the GOAT, praising its direct, sharp, and "human" personality”

来自 Reddit · 英文 · 2026-09-11
查看模型详情 →
#2 Claude Fable 5.1n=35 62.4

“Fable 5.1 is still my go to”

来自 Reddit · 英文 · 2026-09-13

“chewed through my usage for a few turns for sure-- honestly it was pretty nuts”

来自 Reddit · 英文 · 2026-09-12
查看模型详情 →
#3 GPT-5.6 Lunan=11 61.1

“I'm using a lot of gpt-5.6-Luna”

来自 Hacker News · 英文 · 2026-09-12

“right now just gpt Luna is enough I swear”

来自 Reddit · 英文 · 2026-09-12
查看模型详情 →
#1 GLM 5.2n=14 73.2

“it's just that one—like GLM 5.2—charges me $3, whereas DeepSeek v4-flash-0731 did the same job for a ridiculously lower price.”

来自 Reddit · 英文 · 2026-09-12

“glm5.2 连你们所谓的炼炸的 dspro 都不过”

来自 知乎 · 2026-09-10
查看模型详情 →
#2 Gemma 4 31Bn=10 71.9

“gemma4:31b”

来自 Reddit · 英文 · 2026-09-13

“I still use Gemma 4 31B for technical writing. With a tiny bit of prompting, it's a breath of fresh air compared to the junk that Claude, GPT-5.6-sol, and Qwen3.8 27B are outputting these days.”

来自 Reddit · 英文 · 2026-09-09
查看模型详情 →
#3 GLM 5.3 Flashn=66 68.2

“glm flash seems like it doesnt use the memory context effectively”

来自 Reddit · 英文 · 2026-09-13

“around the same level as 5.3-flash in my day to day”

来自 Reddit · 英文 · 2026-09-13
查看模型详情 →