← All models

Gemini · Last 7 days

Gemini 3.8 Flash

0–100 · higher means more positive community experience, not benchmark performance.

Updated 2026-09-05 14:21 UTC · 2026-08-29 – 2026-09-05 UTC
Community score43.5/100140 comments in 7 days

Community reviews

Selected comments from the last 7 days. The balance of excerpts does not represent the share of positive reviews.

6 selected excerpts

PositiveCoding

“DeepSWE v1.1, 3.8 Flash scored 73.7%, an improvement of over 8 percentage points from 3.7 Flash's 65.3%, with Ars Technica reporting it has topped the benchmark. On Terminal-bench 2.1, which focuses on agent-side coding, it achieved 89.4%, surpassing a range of flagship models.”

ZhihuzhMachine translated
Show original

DeepSWE v1.1,3.8 Flash 得分 73.7%,较 3.7 Flash 的 65.3% 提升超过 8 个百分点,Ars Technica 报道称它已登顶该榜单。专注智能体终端编码的 Terminal-bench 2.1 上,它拿到 89.4%,超过了一众旗舰模型。

PositiveCoding

“3.8 flash is definitely good for coding”

ZhihuzhMachine translated
Show original

编程肯定3.8flash

MixedCoding

“gemini 3.8 flash is genuinely excellent at coding. It's the instruction following and making wrong assumptions that suck and the annoying thing where you give it a plan and it comes back with an inferior plan. But one thing i've noticed if”

NegativeCoding

“3.8 flash is benchmaxxed to hell, garbage model”

PositiveCoding

“Based on my experience, its improvement in coding is quite remarkable. Comparing 3.6 -> 3.7 from unusable to usable, I personally think 3.7 -> 3.8 can be considered from usable to good”

ZhihuzhMachine translated
Show original

基于我的使用体验,它在 coding 方面的提升是非常不错的。对比 3.6 -> 3.7 从不可用到可用,其实我个人觉得 3.7 -> 3.8 可以算是 可用到好用了

NeutralCoding

“Deep SWE is single-task short programming, while Terminal Bench 4 is long-task programming averaging five to six hours. Flash model, Terminal Bench 4 basically won't be very high”

ZhihuzhMachine translated
Show original

Deep SWE 是单任务短编程,而 Terminal Bench 4 是平均五六个小时的长任务编程。Flash 模型,Terminal Bench 4 基本不会很高