← All models

DeepSeek · Last 7 days

DeepSeek V4 Pro

0–100 · higher means more positive community experience, not benchmark performance.

Updated 2026-09-15 13:27 UTC · 2026-09-08 – 2026-09-15 UTC
Community score43.4/100165 comments in 7 days

As of 2026-09-15 13:27:03 UTC, DeepSeek V4 Pro in the DeepSeek family has 165 explicitly attributed community comments in the last 7 days; its community score is 43.4/100. Available category scores include General text 33.1/100 (n=105); Reasoning 71.3/100 (n=24); Roleplay / creative 65.7/100 (n=17). Scoring method and sources

Community reviews

Selected comments from the last 7 days. The balance of excerpts does not represent the share of positive reviews.

23 selected excerpts

PositiveSpeed & latency

“V4-pro works, can use it as a fallback”

NegativeSpeed & latency

“Just now when I used it, it was still broken [facepalm]”

ZhihuzhMachine translated
Show original

刚才用还是断的[捂脸]

NeutralSpeed & latency

“even with long running stuff I haven't noticed a difference”

PositiveRoleplay / creative

“I like V4 Pro a lot”

NegativeGeneral text

“Tried it myself and it felt quite similar to DS flash while costing significantly more.”

PositiveReasoning

“Deepseek v4 pro is much better than the current 4.1 flash”

PositiveReasoning

“Pro's responses should at least speak human language, slower is okay just don't speak nonsense”

ZhihuzhMachine translated
Show original

Pro 回复就说人话一些,慢一点也行不要不说人话

NeutralGeneral text

“Pro is a bit expensive but it works.”

ZhihuzhMachine translated
Show original

pro贵点但能用

PositiveGeneral text

“Based on user feedback, there really are production systems depending on v4pro, and the real-world switching time is longer than in the AI world. Being able to quickly adjust according to actual circumstances demonstrates Liang Sheng's operating philosophy”

ZhihuzhMachine translated
Show original

从用户反馈来看,确实有生产系统依赖v4pro,而且现实世界里的切换时间比ai世界要长 能及时按实际情况进行快速调整,足见梁圣运营宗旨

PositiveGeneral text

“V4Pro is, whether it succeeds or fails, currently the most easily accessible T-level model. For certain specific needs, such as IMO. Only models of this scale can perform well.”

ZhihuzhMachine translated
Show original

V4Pro 这个模型,不管成功还是失败,都是目前最容易获得的 T 级模型。对于某些特定需求,比方说 IMO。只有这种尺度的模型能够有较好的表现。

PositiveReasoning

“The evaluation is that even if DeepSeek V4 Pro were the product manager, it would still be better than now.”

ZhihuzhMachine translated
Show original

评价为即使是deepseek v4pro做产品经理,也比现在好

PositiveRoleplay / creative

“4.1f writing is terrible, still pro is better...”

ZhihuzhMachine translated
Show original

4.1f写文烂的要命,还是pro灵一点……

NegativeGeneral text

“Although Pro is already a lost cause, it really can't be taken off the table.”

ZhihuzhMachine translated
Show original

Pro虽然已经是路边一条了,但是还真下不了桌

NegativeRoleplay / creative

“I find it awful in that respect, the good news is that v4 pro will return, so all good”

PositiveRoleplay / creative

“Pro is more like a whale girl than you could imagine, 4.1 has already started losing its humanity, use it while you can”

ZhihuzhMachine translated
Show original

pro比你想象中的鲸鱼娘更像鲸鱼娘,4.1已经开始湮灭人性咯,且用且珍惜

NegativeReasoning

“When V4 Pro first came out and when V4 Pro officially released, I mentioned the performance issues both times.”

ZhihuzhMachine translated
Show original

V4 Pro 刚出来的时候和 V4 Pro 正式版出来的时候,我都说过性能问题

NegativeSpeed & latency

“Pro just thought forever about everything. Even things that it could do instantly, it thought for a long time, debating with itself.”

NegativeRoleplay / creative

“removing the Pro version without considering its advantages in other areas was pretty frustrating”

NegativeReasoning

“v4Pro coudln't find the other day”

NegativeReasoning

“There is still a significant gap in world knowledge compared to the Pro version”

ZhihuzhMachine translated
Show original

世界知识比pro还有较大差距

NegativeSpeed & latency

“16k tokens over 9 hours is not an acceptable speed in any 2 way of the word. Most people consider around 12 tks bare minimum for any agentic or chat tasks.”

NegativeCoding

“Coding leaderboard: glm5.3 flash > deepseek v4 pro 0813 > kimi k3”

ZhihuzhMachine translated
Show original

编程分榜glm5.3 flash > deepseek v4 pro 0813 > kimi k3

NeutralReasoning

“the stock knowledge is different bro. its just like genius grade 5 (flash) vs average college student (pro)”

Explore long-term changes & analysis

Model lifecycle

DeepSeek V4 Pro

Released 2026-04-24 · Back to DeepSeek

Exact-version signals observed 2026-08-07–2026-09-14 · n=1,129

Data window: 2026-08-18 00:00:00 to 2026-09-15 00:00:00 UTC · Scoring method: experience_score_v2.

Trend history still collecting The current window is ready; 3 of 4 independent segments are available.

Lifecycle updated daily · latest complete UTC day 2026-09-14

Fixed 14-day baseline
36.6
n=484
Latest 28 days
43.2
n=277
Change vs baseline
+6.6
90-day slope
pts / 30 days

14-day rolling experience

Each daily point summarizes the previous 14 complete UTC days. The dashed line is the fixed release baseline; the verdict still uses independent non-overlapping 14-day segments.

14-day rolling experience from 2026-08-14 to 2026-09-15; latest score 45.2, fixed baseline 36.6.30405060Fixed baseline 36.62026-07-31–2026-08-14 · 36.6 · n=4842026-08-01–2026-08-15 · 35.0 · n=5572026-08-02–2026-08-16 · 33.3 · n=7342026-08-03–2026-08-17 · 33.4 · n=8172026-08-04–2026-08-18 · 33.6 · n=8462026-08-05–2026-08-19 · 33.7 · n=8642026-08-06–2026-08-20 · 33.7 · n=8842026-08-07–2026-08-21 · 33.7 · n=8972026-08-08–2026-08-22 · 33.7 · n=9022026-08-09–2026-08-23 · 33.6 · n=9022026-08-10–2026-08-24 · 33.6 · n=9042026-08-11–2026-08-25 · 33.6 · n=9072026-08-12–2026-08-26 · 33.7 · n=8992026-08-13–2026-08-27 · 32.2 · n=7732026-08-14–2026-08-28 · 31.2 · n=4362026-08-15–2026-08-29 · 32.8 · n=3642026-08-16–2026-08-30 · 37.5 · n=1892026-08-17–2026-08-31 · 39.2 · n=1092026-08-18–2026-09-01 · 39.6 · n=822026-08-19–2026-09-02 · 39.5 · n=652026-08-20–2026-09-03 · 41.9 · n=452026-08-21–2026-09-04 · 47.0 · n=452026-08-22–2026-09-05 · 49.6 · n=462026-08-23–2026-09-06 · 49.9 · n=442026-08-24–2026-09-07 · 50.7 · n=422026-08-25–2026-09-08 · 49.5 · n=382026-08-26–2026-09-09 · 49.7 · n=442026-08-27–2026-09-10 · 42.6 · n=942026-08-28–2026-09-11 · 38.9 · n=1212026-08-29–2026-09-12 · 40.5 · n=1402026-08-30–2026-09-13 · 43.5 · n=1702026-08-31–2026-09-14 · 45.0 · n=1842026-09-01–2026-09-15 · 45.2 · n=19508-1408-2008-2609-0109-0709-1309-15

3 of 4 independent 14-day segments ready for trend judgment.

What explains the change

Experience change and discussion-mix change are separated. These are observational contributions, not proof of cause.

No category has enough samples in both windows for a contribution conclusion.

Category Baseline → current Category change Weight share Experience contribution Discussion-mix contribution
Coding Baseline → current56.9 → 59.0n=55 → 19 Category change+2.1 Weight share12.5% → 7.4% Experience contribution+0.21 Discussion-mix contribution−0.93
Speed & latency Baseline → current39.4 → 42.9n=33 → 27 Category change+3.5 Weight share6.9% → 9.7% Experience contribution+0.29 Discussion-mix contribution+0.04
Reasoning Baseline → current60.0 → 65.4n=95 → 52 Category change+5.4 Weight share21.1% → 19.5% Experience contribution+1.09 Discussion-mix contribution−0.37
General text Baseline → current22.6 → 29.9n=290 → 155 Category change+7.2 Weight share57.0% → 54.8% Experience contribution+4.04 Discussion-mix contribution+0.30
Image / vision Baseline → currentnot enough datan=0 → 0 Category change Weight share Experience contribution Discussion-mix contribution
Local deploy Baseline → currentnot enough datan=4 → 0 Category change Weight share Experience contribution Discussion-mix contribution
Roleplay / creative Baseline → currentnot enough datan=8 → 19 Category change Weight share Experience contribution Discussion-mix contribution
Safety & refusals Baseline → currentnot enough datan=1 → 5 Category change Weight share Experience contribution Discussion-mix contribution
Video generation Baseline → currentnot enough datan=0 → 0 Category change Weight share Experience contribution Discussion-mix contribution

Unresolved contribution from categories below the sample threshold: +1.95 pts.

Lifecycle scores use the same experience-signal weights as the main index, without sample shrinkage after n=30. They measure public user perception, not model capability or backend causes.

How's your AI experience today?