← All models

Gemini · Last 7 days

Gemini 3.8 Flash

0–100 · higher means more positive community experience, not benchmark performance.

Updated 2026-09-05 14:21 UTC · 2026-08-29 – 2026-09-05 UTC
Community score43.5/100140 comments in 7 days

Community reviews

Selected comments from the last 7 days. The balance of excerpts does not represent the share of positive reviews.

29 selected excerpts

MixedSpeed & latency

“I'd say it's about equal to luna, but FAST. And terra is definitely more intelligent. Gemini 3.8 flash still has that old gemini problem of often just going for the quick surface solution.”

PositiveReasoning

“It can disentangle and understand a 100,000-word mystery novel with a chaotic timeline and narrative tricks, something none of the other models I tried could do.”

ZhihuzhMachine translated
Show original

它能梳离明白一本时间线混乱,叙述性诡计的十万字推理小说,我试的其他模型都做不到

PositiveGeneral text

“Gemini 3.8 flash is actually pretty good now.”

PositiveReasoning

“I did gem 3.8 flash on medium I feel its the perfect one for that because sometimes it can over think”

NegativeSpeed & latency

“It still drools or slacks off when the process gets long, so I gave up on it early and use it for coding instead.”

ZhihuzhMachine translated
Show original

依旧流程一长不是流口水就是偷懒,及早放弃用来coding

PositiveGeneral text

“I think 3.7 and 3.8 are a significant improvement over 3.1.”

ZhihuzhMachine translated
Show original

我觉得3.7 3.8 比3.1要提高不少

NegativeGeneral text

“This one doesn't perform as well as GPT.”

ZhihuzhMachine translated
Show original

这一比gpt不太行

NegativeGeneral text

“Same model, one version of flash, but the pro suddenly became very cold and formal. Also doesn't like speaking in a 1, 2, 3 style anymore. Analogies have all disappeared.”

ZhihuzhMachine translated
Show original

一样的模型,一个号的flash,pro突然变得非常高冷,说话一板一眼的了。而且说话也不爱1,2,3这种方式。类比全部消失。

NegativeSafety & refusals

“Gemini confidently ignores instructions. I thought things would get better with 3.8, but somehow it feels even worse. Even when I explicitly tell it something like "just tell me", it goes straight into editing files.”

PositiveCoding

“DeepSWE v1.1, 3.8 Flash scored 73.7%, an improvement of over 8 percentage points from 3.7 Flash's 65.3%, with Ars Technica reporting it has topped the benchmark. On Terminal-bench 2.1, which focuses on agent-side coding, it achieved 89.4%, surpassing a range of flagship models.”

ZhihuzhMachine translated
Show original

DeepSWE v1.1,3.8 Flash 得分 73.7%,较 3.7 Flash 的 65.3% 提升超过 8 个百分点,Ars Technica 报道称它已登顶该榜单。专注智能体终端编码的 Terminal-bench 2.1 上,它拿到 89.4%,超过了一众旗舰模型。

PositiveGeneral text

“Luna sucks. Use Gemini 3.8 flash ffs”

MixedGeneral text

“You say it's not good? It has a few yuan 18-month subscription on the second-hand market, I've been using it for at least 3 weeks and already gotten my money's worth, its capability is definitely above the value of 10 RMB. You say it's good? 30% of the time it needs another model to wipe its butt.”

ZhihuzhMachine translated
Show original

你说它不行吧,它有海鲜市场几块钱18个月的订阅,我至少撑了3周了已经回本了,它的水平肯定高于10块人民币提供的价值。你说它行吧,30%的情况下需要别的模型擦腚。

PositiveCoding

“3.8 flash is definitely good for coding”

ZhihuzhMachine translated
Show original

编程肯定3.8flash

MixedCoding

“gemini 3.8 flash is genuinely excellent at coding. It's the instruction following and making wrong assumptions that suck and the annoying thing where you give it a plan and it comes back with an inferior plan. But one thing i've noticed if”

PositiveSpeed & latency

“Gemini 3.8 Flash (which is an _excellent_ model especially for its price and TPS!!”

NegativeGeneral text

“Not as good as Doubao”

ZhihuzhMachine translated
Show original

不如豆包

PositiveSpeed & latency

“Gemini 3.8 flash is amazing for a small model”

NegativeReasoning

“Son raisonnement est merdique comparé à gemma 4 e2b en qat”

NegativeCoding

“3.8 flash is benchmaxxed to hell, garbage model”

NegativeReasoning

“Still top-tier at benchmark grinding, top-tier at scores, actual performance is garbage”

ZhihuzhMachine translated
Show original

依旧刷题一流,跑分一流,实际一坨。

NegativeSpeed & latency

“gemini 3.8 flash spent 2 minutes writing more than two screens worth, I couldn't be bothered to read what its main points were”

ZhihuzhMachine translated
Show original

gemini 3.8 flash 花了 2 分钟写了两屏幕多,我都懒得看它写的重点是什么

NegativeSpeed & latency

“Besides being fast, it doesn't seem to have other advantages”

ZhihuzhMachine translated
Show original

除了快好像没别的优点了

NegativeReasoning

“Ignoring prompts, agents.md and other file requirements, being smartass, resulting in output that looks pretty but is actually a pile of dog shit”

ZhihuzhMachine translated
Show original

无视提示词,agents.md,等文件要求 自作聪明,导致结果看起来很漂亮实际上就是一坨狗屎

NeutralGeneral text

“Flash's peers are Sonnet/Terra; comparing it with Opus/Sol is really giving Google too much credit”

ZhihuzhMachine translated
Show original

Flash的同级别对手是Sonnet/Terra,和Opus/Sol比还真是高看谷歌啊

PositiveCoding

“Based on my experience, its improvement in coding is quite remarkable. Comparing 3.6 -> 3.7 from unusable to usable, I personally think 3.7 -> 3.8 can be considered from usable to good”

ZhihuzhMachine translated
Show original

基于我的使用体验,它在 coding 方面的提升是非常不错的。对比 3.6 -> 3.7 从不可用到可用,其实我个人觉得 3.7 -> 3.8 可以算是 可用到好用了

PositiveReasoning

“Requires you to put effort into task decomposition and prompts, but compared to 3.7 which couldn't even execute well-made plans, 3.8 is much stronger”

ZhihuzhMachine translated
Show original

需要你在任务分解和提示词上费功夫,但是相比3.7连写好的计划都无法实现来说,3.8还是强很多

PositiveRoleplay / creative

“Don't know about the rest, but RP enthusiasts will love it, feels very alive”

ZhihuzhMachine translated
Show original

其他不知道,但rp🦌客狂喜,活人感极强

NeutralCoding

“Deep SWE is single-task short programming, while Terminal Bench 4 is long-task programming averaging five to six hours. Flash model, Terminal Bench 4 basically won't be very high”

ZhihuzhMachine translated
Show original

Deep SWE 是单任务短编程,而 Terminal Bench 4 是平均五六个小时的长任务编程。Flash 模型,Terminal Bench 4 基本不会很高

NeutralSpeed & latency

“gemini's output speed is 300t/s”

ZhihuzhMachine translated
Show original

gemini的输出速度是300t/s