Current ranking
Speed & latency
For the 7-day window ending 2026-09-13 23:17:02 UTC, user-experience leaders in Speed & latency: closed models — GPT Image 2.5 (73.6/100, n=8); open models — Qwen3.6 35B A3B (70.1/100, n=18).
#1 GPT Image 2.5n=8 73.6▾
Open model detail →“还有,现在在 codex 里面出的图,也是直接透明背景的了。你告诉它出图的时候做成透明背景就行。这一点我要着重说一下,可能很多做游戏角色的人都是这样子的:先出一个游戏的角色原型图,但是因为要把它周围的图抠掉,所以要先填充一个纯绿色的图,整个图出完之后,再把绿色抠掉。如果你只抠一个还好,如果你大批量地抠,这真的是灾难性的工作。便用 AI 来抠,也有很多残留。这些问题都解决了。(2.0 就解决了)”
via Zhihu · Chinese · 2026-09-09
#2 Gemini 3.8 Flashn=9 65.3▾
“Gemini 3.8 flash à 1200 token par seconde sur antigravity.”
via Reddit · 2026-09-13
Open model detail →“That's Gemini 3.8 Flash right now. I have it tearing shit apart because it's so fast that I can afford the time to be like "hey uhh... Check this thing out for me."”
via Reddit · 2026-09-10
#3 GPT-5.6 Lunan=5 64.4▾
“Ds 4 flash had to think twice as much as gpt 5.6 luna for simple tasks”
via Reddit · 2026-09-10
Open model detail →“小尺寸廉价模型能打的,基本就ds v4flash, glm5.3 flash跟gpt luna这仨选择吧”
via Zhihu · Chinese · 2026-09-10
#1 Qwen3.6 35B A3Bn=18 70.1▾
“larger model performed a little better, albeit a little slow (37-40 tps)”
via Reddit · 2026-09-13
Open model detail →“qwen35b a3b试试vulkan,比rocm快”
via Zhihu · Chinese · 2026-09-13
#2 DeepSeek V4 Flashn=70 69.7▾
“the speed and the clarity of communication is simply unmatched by any so-called "frontier model"”
via Reddit · 2026-09-12
Open model detail →“Locally I use Deepseek Flash, Qwen3.8-Next-Flash, or Gemma 4 depending on how slowly I want my slop to generate.”
via Hacker News · 2026-09-12
#3 Qwen3.8 Flash Nextn=64 69.2▾
“Its getting 750 t/s prefill and 40 t/s output even at 200k contexts”
via Reddit · 2026-09-13
Open model detail →“I was able to get 6t/s with qwen3.8 flash next with q4_k_xl using ssd offloading for the inactive layers”
via Reddit · 2026-09-13