← All models

Qwen · Last 7 days

Qwen3.8 Flash Next

0–100 · higher means more positive community experience, not benchmark performance.

Updated 2026-09-05 14:21 UTC · 2026-08-29 – 2026-09-05 UTC
Community score61.0/100107 comments in 7 days

Community reviews

Selected comments from the last 7 days. The balance of excerpts does not represent the share of positive reviews.

14 selected excerpts

NeutralSpeed & latency

“Prompt prosessing changed a lot at lower pcie bandwidth while token generation was influenced a lot less”

PositiveReasoning

“This is a frontier-class model”

NeutralLocal deploy

“Qwen3.8 flash next you can probably do a Q4 or maybe even Q5, but you'll actually need 64 GB system RAM because of the 51B n-gram in that model.”

PositiveReasoning

“It reasons well, it covers 90% of what frontier models used to do for me.”

PositiveReasoning

“king of MOEs right now”

NegativeGeneral text

“Qwen3.8 flash next's implementation is very unstable in many inference engine. Give it more time.”

PositiveGeneral text

“Switched to it from DS4F for RAG summarizer, quality is definitely up, no garble. vLLM on sparks, hybrid NVFP4+FP8 checkpoint.”

NegativeReasoning

“qwen3.8-flash-next's thinking is particularly long at xhigh level, way longer than glm-5.3-flash's max level”

RedNotezhMachine translated
Show original

qwen3.8-flash-next在xhigh档位下思考特别长,比glm-5.3-flash的max档位长太多了

PositiveRoleplay / creative

“doing a really good job writing creative stories”

PositiveLocal deploy

“Qwen 3.8 Flash Next is a big step in this direction”

PositiveCoding

“pretty impressed with what it can do with novel prompts”

NegativeGeneral text

“using nvfp4 radix quant the quality suffered compared to deepseek”

NegativeGeneral text

“I have been working with Qwen3.8 Flash Next, and while the level of speech is not amazing, I haven't seen anything like this in my custom harness.”

PositiveLocal deploy

“fits on a 256GB system in Q8”