Model recommendations

Which model feels best right now, by task — ranked by real user experience over the last 7 days.

Same 7-day score · updated 2026-09-11 16:05 UTC

Top 3 specific models in each category. A model appears when its category sample size is at least 5. Rank movement compares the previous published snapshot. Open a ranked model to read matching user quotes when available; rows without quotes link straight to the model page.

General text

#1 Claude Opus 4.6n=8 72.7

“massive love-fest for Opus 4.6. Y'all think it's the GOAT, praising its direct, sharp, and "human" personality”

via Reddit · 2026-09-11

“Opus 4.6 with medium effort implementing plans written by Fable 5.1”

via Reddit · 2026-09-10
Open model detail →
#2 Gemini 3.8 Flashn=18 59.9

“我觉得大概率是降智了,所以我也没有大规模的给人推荐,Gemini 3.8 Flash。 因为上次3.7的时候,我一开始用的挺好,然后就给别人推荐了一下,结果用了不到一周的时间,就突然之间幻觉爆棚,我就放弃了。[捂脸]”

via Zhihu · Chinese · 2026-09-11

“3.8智力比以前明显感觉到提升,但是经常自作主张和任务不完成或者完成的一半就说自己完成了。最傻的是他自己的总结计划里都没说自己完成了还告诉我他全部任务我已经做完”

via Zhihu · Chinese · 2026-09-10
Open model detail →
#3 GPT-5.5n=9 57.8

“GPT-5.5 at the same reasoning level does not have this. Same account, same standard chat, same prompt, minutes apart: 5.6 Sol fails every time, 5.5 completes and actually runs the search.”

via Reddit · 2026-09-09

“A good plan makes 5.5 sufficient; a better model may reduce how much planning and supervision you need.”

via Reddit · 2026-09-08
Open model detail →

Coding

#1 Claude Sonnet 5n=7 73.1
#2 Claude Fable 5n=75 68.0

“Fable did pretty well on that aspect and also the UI was decent”

via Reddit · 2026-09-11

“Fable for coding, Sol for code review”

via Reddit · 2026-09-10
Open model detail →
#3 Claude Fable 5.1n=38 63.4

“exactly, fable wins this one”

via Reddit · 2026-09-11

“way better, plays way better”

via Reddit · 2026-09-11
Open model detail →

Reasoning

#1 Claude Fable 5n=73 81.3

“Fable is very very good at long context through repeated compactions. With a proper handoff file it follows along very very well.”

via Reddit · 2026-09-11

“Alpöge himself presented a counterexample to the Jacobian conjecture on July, found with Claude Fable”

via Hacker News · 2026-09-11
Open model detail →
#2↑1 GPT-6 Astran=268 81.0

“Same thing happened to me. When I asked it to stop, it started engineering an elaborate rollback process, talking about what would happen if two people tried rolling back simultaneously, etc.”

via Reddit · 2026-09-11

“Astra is way more capable and seems more like a genuine collaborator that Opus pretends to be and yeah it also gets things done a bit better”

via Reddit · 2026-09-11
Open model detail →
#3↓1 Claude Fable 5.1n=38 80.6

“Fable understood the assignment”

via Reddit · 2026-09-11

“Fable seems to forget about things”

via Reddit · 2026-09-10
Open model detail →

Image / vision

#1 GPT-5.4 Image 2n=13 66.6

“generates more realistic images than gemini”

via Hacker News · 2026-09-09

“明显是2的意境更好”

via Zhihu · Chinese · 2026-09-09
Open model detail →
#2 GPT Image 2.5n=50 66.3

“连续编辑值得试,它让已经接近可用的图片更方便地继续修改”

via Zhihu · Chinese · 2026-09-10

“动作理解相当厉害,手指依旧困难,定点改图的能力极大提升”

via Zhihu · Chinese · 2026-09-10
Open model detail →
#3 GPT-6 Astran=14 65.6

“Astra is a massive step up for me because it understand images (view image tool) muuuuuch better than SOL”

via Reddit · 2026-09-09

“怎么最后出来的图像《模拟城市》一样”

via Zhihu · Chinese · 2026-09-09
Open model detail →

Roleplay / creative

#1 GPT-6 Astran=14 47.0

“写的很好啊,超过大多数人了”

via Zhihu · Chinese · 2026-09-05

“这个故事写的是真有点感觉,感到震撼了”

via Zhihu · Chinese · 2026-09-05
Open model detail →
#2 GPT-5.6 Soln=5 38.9

“yes i described the witcher to it but not directly so it did a pretty good job knowing who i was on about”

via Reddit · 2026-09-09

“I think medium and high settings write better, based on my experience.”

via Reddit · 2026-09-09
Open model detail →
#3 Grok 4.6n=5 36.1

“Newer ones have every character speak in complete, helpful, on-topic sentences — because that's what the assistant training rewards. Everyone sounds like the same person wearing different hats.”

via Reddit · 2026-09-08

“4 and 4.1 were amazing. 4.6 is useless. Purely a downgrade.”

via Reddit · 2026-09-08
Open model detail →

Speed & latency

#1 GPT Image 2.5n=7 72.4

“还有,现在在 codex 里面出的图,也是直接透明背景的了。你告诉它出图的时候做成透明背景就行。这一点我要着重说一下,可能很多做游戏角色的人都是这样子的:先出一个游戏的角色原型图,但是因为要把它周围的图抠掉,所以要先填充一个纯绿色的图,整个图出完之后,再把绿色抠掉。如果你只抠一个还好,如果你大批量地抠,这真的是灾难性的工作。便用 AI 来抠,也有很多残留。这些问题都解决了。(2.0 就解决了)”

via Zhihu · Chinese · 2026-09-09
Open model detail →
#2 GPT-5.6 Lunan=6 68.2

“Ds 4 flash had to think twice as much as gpt 5.6 luna for simple tasks”

via Reddit · 2026-09-10

“小尺寸廉价模型能打的,基本就ds v4flash, glm5.3 flash跟gpt luna这仨选择吧”

via Zhihu · Chinese · 2026-09-10
Open model detail →
#3 Gemini 3.8 Flashn=13 63.0

“That's Gemini 3.8 Flash right now. I have it tearing shit apart because it's so fast that I can afford the time to be like "hey uhh... Check this thing out for me."”

via Reddit · 2026-09-10

“cant generate in ai studio answer after 5-7min xd”

via Reddit · 2026-09-08
Open model detail →

Top 3 specific models in each category. A model appears when its category sample size is at least 5. Rank movement compares the previous published snapshot. Open a ranked model to read matching user quotes when available; rows without quotes link straight to the model page.

General text

#1 GLM 5.2n=17 75.8

“glm5.2 连你们所谓的炼炸的 dspro 都不过”

via Zhihu · Chinese · 2026-09-10

“As of today, clearly Z.Ai. Both its Pro and Flash models outperform DS's, and the subscription plans are likely the same price or cheaper”

via Reddit · 2026-09-09
Open model detail →
#2 GLM 5.3 Flashn=59 74.5

“Glm 5.3 flash is equivalent to 5.6 terra , not sol, sol and 5.3 is equivalent, sol is kinda a lot better Astra is a lot lot lot better”

via Reddit · 2026-09-11

“体感上不如GLM 5.3 FLASH,就是快是真快,但是没写对”

via Zhihu · Chinese · 2026-09-11
Open model detail →
#3 DeepSeek V4 Flashn=72 62.6

“pro v4 was great, this switch to flash has stripped a lot of nuance from it.. It's become very blunt, direct and even aggressive. The analyis and ideas are actually strong though, better than before, very good on that front, but socially it”

via Reddit · 2026-09-10

“直接连v4好像连不上了”

via Zhihu · Chinese · 2026-09-10
Open model detail →

Coding

#1 GLM 5.3 Flashn=24 72.1

“return to glm 5.3 flash for instaprogress”

via Reddit · 2026-09-10

“I was writing low level code and the 0731 was a lot better than GLM 5.3 Flash”

via Reddit · 2026-09-10
Open model detail →
#2 Qwen3.8 27Bn=40 68.9

“can confirm that Qwen3.8 27b on a 3090 (24gb) is a damn fine cup of coffee for coding and general use”

via Reddit · 2026-09-10

“Qwen 3.8 has been working on the game's entire framework like a tireless little ant”

via Reddit · 2026-09-09
Open model detail →
#3 DeepSeek V4 Flashn=13 67.1

“They handle most everyday coding-agent tasks fine, and you can run them usage-based or on a flat plan”

via Reddit · 2026-09-11

“it should have no problem adapting to your codebase. Get yourself an API key, install "dsh", put in the key, and get coding with deepseek-v4-flash”

via Reddit · 2026-09-10
Open model detail →

Reasoning

#1 Qwen3.8 27Bn=55 80.5

“going from 3.6-27B to 3.8.27B is that it works much better to solve problems over time, figuring out how to gather data, trying multiple approaches, not getting down an endless rabbit hole, etc.”

via Reddit · 2026-09-11

“I'm Qwen3.8 27B to some times overthink itself to the wrong answer, will have to compare”

via Reddit · 2026-09-10
Open model detail →
#2 GLM 5.3 Flashn=32 79.3

“I have a Synthetic subscription and really like it for GLM. I get tons of usage out of 5.3 Flash and it's stupid capable.”

via Reddit · 2026-09-11

“你说的挺对的,我用着也是glm5.3 flash强”

via RedNote · Chinese · 2026-09-10
Open model detail →
#3 Qwen3.8 Flashn=10 77.2

“The model is very accurate and, except for speed, far surpasses 27b.”

via Reddit · 2026-09-10

“在dsh上的思考深度与全面性,令我惊叹”

via Zhihu · Chinese · 2026-09-08
Open model detail →

Speed & latency

#1 DeepSeek V4 Flashn=68 72.3

“个人体感就是做视觉方面的还是很差,一个任务做了快1小时,没有任何反应,用Cursor 几分钟做好了”

via Zhihu · Chinese · 2026-09-11

“快,但是老要返工,就很难评。。。”

via Zhihu · Chinese · 2026-09-11
Open model detail →
#2 Qwen3.6 35B A3Bn=12 67.7

“能有700多token的输入速度,和30多Token的输出速度”

via Zhihu · Chinese · 2026-09-11

“it will run about 35 tok/s if you set it up right with enough context for most things”

via Reddit · 2026-09-06
Open model detail →
#3↑1 Qwen3.8 Flash Nextn=57 67.6

“I even managed to get IQ\_4\_XS Qwen3.8 Flash Next running at \~15tok/s and 65k context which is wild for a 100GB model file (admittedly it ate up every single MB of my 64GB of RAM and even my SSD was at a solid 50% utilization at points so”

via Reddit · 2026-09-11

“25-30 tokens per second with this config - 5090 & 256GB DDR5”

via Reddit · 2026-09-11
Open model detail →

Local deploy

#1 Qwen3.8 27Bn=36 69.1

“great models like Qwen3.8-27B that I can run at home with Open Weights under the Apache 2.0 for free!”

via Reddit · 2026-09-10

“can fit comfortably in 16g vram with 100k+ context plus MTP or vision projector”

via Reddit · 2026-09-10
Open model detail →
#2↑1 Qwen3.6 35B A3Bn=5 66.8

“You can run qwen3.6 35B with an 8gb card at 35-40 tok/s at about 5.6gb on the card for a XS Q4 but it's a tight fit when you add another 1gb for the mmproj vision.”

via Reddit · 2026-09-05
Open model detail →
#3↓1 GLM 5.3 Flashn=9 65.5

“GLM 5.3 Flash is a clear winner for 128gb vram, for me.”

via Reddit · 2026-09-10

“GLM 5.3 Flash at Q4”

via Reddit · 2026-09-10
Open model detail →