DeepSeek Price Hike Sparks Exodus as MiMo V2.5 Emerges as Developers' Last Affordable Option

Flash tier jumps 4–5× on August 17, OpenCode Go slashes DeepSeek quota 94 %, and Qwen3.8-27B runs Opus 4.6-class tasks on a single 24 GB GPU

Hacker News 2 · Reddit 8 · Zhihu 6 964 covered discussions 5 source-linked evidence passages

The day in brief

DeepSeek's new time-of-day pricing kicks in, doubling costs during peak hours and raising V4 Flash 4–5× for light users — users call it the empire's collapse and plan migrations to Qwen3.8-27B and GPT Luna

OpenCode Go's DeepSeek Flash allocation crashes 94 % from 63,300 to 3,800 tokens in five hours; MiMo V2.5 becomes the only plan with enough monthly tokens for moderate development work

Qwen3.8-27B posts benchmark scores neck-and-neck with DeepSeek V4 and GPT-5.6 Luna, running at 30–70 tok/s on consumer GPUs from a 3090 to a 7900 XTX; Simon Willison flags overthinking that reasoning budget controls can fix

Claude faces a dual crisis: 20× Pro users burn through weekly limits in hours, authentication goes down, and Opus 5/Fable 5 degrade with odd jargon, casual tone, and persistent disclaimers

GPT-5.6 Sol earns praise as OpenAI's best vision model, outreasoning Claude-produced UIs on visual tasks without tool assistance

01

Product and platform changes

01

DeepSeek's New Peak/Off-Peak Pricing Doubles Costs; Light Users Face 4–5× V4 Flash Hike

What Happened

DeepSeek's new pricing went live at midnight on August 17, replacing flat-rate billing with a time-of-day tiered structure.

Peak hours — 9:00–12:00 and 14:00–18:00 daily — now cost twice as much as off-peak periods.

Light users who previously spent 2.5 months consuming 1.52 trillion tokens face monthly bills exceeding 4,000 yuan under the new structure.

A user calculated that the new pricing could fund a Codex x5 Pro Plan with room to spare, or a x20 Pro Plan including Sol with some extra spend.

V4 Pro exited its trial phase and entered commercial availability with the Responses API and Codex support at a higher price tier.

Long-context sessions burn through tokens rapidly; debugging a system prompt can easily generate hundreds of thousands of input tokens per turn.

Users have begun migrating housekeeping tasks to Qwen 3.8 27B, calling DeepSeek's price hike a dug moat.

Community Reaction

After venting, users did the math on competitors and found Flash still cheapest at its tier — but could not locate a cheaper alternative.

Users criticized DeepSeek for the classic move of building user habits with low prices then gradually raising them.

Heavy users complained about the absence of volume discounts; the bigger the usage, the larger the absolute price increase.

Some users said they would stick with DeepSeek only for already-deployed features and switch everything else to cheaper options.

Expectations are that most casual users will be driven away, with prices only dropping once compute capacity catches up.

Practical Takeaway

New pricing demands a full cost-structure reassessment; light-to-moderate users may see monthly spending increase by thousands of yuan.

Peak-hour usage (9–12 and 14–18) now costs double — consider shifting to off-peak windows.

Agent housekeeping tasks are candidates for migration to Qwen 3.8 27B via local deployment.

GPT Luna has already surpassed DeepSeek on cost-performance for certain scenarios.

Use Case

Everyday conversation and simple edits for light non-professional users.

Complex agent chains requiring the Responses API and Codex support.

Budget planning for long-context debugging sessions.

Practical Value

Helps users quickly decide whether to stay with DeepSeek or migrate.

Provides cost comparisons for peak versus off-peak usage.

Identifies task types that Qwen 3.8 27B can absorb.

Community evidence

It's too expensive, even more costly than I expected... As a light-to-medium non-professional user, I used about 152 billion tokens over 2.5 months. Besides DeepSeek, I also frequently use Gemini, ChatGPT web version, and Codex. Previously, spending 20 yuan a day was considered a lot, but this morning after just casual use for a while, asking for modification suggestions on some programs I'd already completed, 36.28 yuan was gone... Regarding the new price calculation, the new DeepSeek expense level is really not worth it in my opinion... Except for some features already deployed on it, I'm planning to switch most of my other usage to other models, and I'm currently researching other more affordable plan options.

02

OpenCode Go Flash Quota Drops 94 % — MiMo V2.5 Only Plan with Enough Monthly Tokens

What Happened

OpenCode Go's Flash plan allocation plummeted from 63,300 tokens to 3,800 tokens within five hours — a 94 % reduction.

Moderate developer usage runs roughly 100 billion tokens per month; at the new allowance V4 Flash would last just seven working days.

V4 Pro's monthly allocation would sustain approximately two working days at peak pricing or four working days off-peak.

A detailed token-economy analysis shows MiMo V2.5 is the only plan exceeding 10 billion monthly tokens at 10.921 billion.

DeepSeek V4 Flash at peak pricing delivers 68 million monthly tokens; GPT-5.6 Luna at the ≤272K tier reaches 52.5 million.

Some users said OpenCode Go was previously so generous they never exhausted it in a month — that is no longer the case.

Kimi K2.7 Code and GLM-5.3 offer only tens of millions of tokens monthly, falling far short.

Community Reaction

Users lament that 'the sky has collapsed' — the best cost-performance DeepSeek package is gone.

Some believe the original V4 Flash allowance was simply too generous and invited today's cuts.

A portion of users are actively researching other, more affordable plan options.

Others argue DeepSeek's raw performance is merely mediocre, trading blows with Gemini.

Practical Takeaway

OpenCode Go's DeepSeek offering can no longer support moderate development usage.

MiMo V2.5 (72,625 tokens per call, 150,376 calls per month) is currently the only sufficiently provisioned option.

GPT Luna (≤272K) at 52.5 million monthly tokens comes in second.

Users may need to purchase multiple plans or migrate across models to cover actual needs.

Use Case

Monthly token demand assessment for moderate developer workflows.

OpenCode Go plan selection decisions.

Cross-model monthly cost-to-token value comparisons.

Practical Value

Helps users quickly assess whether OpenCode Go remains worth using.

Supplies a detailed per-model monthly allowance and supported request-count reference table.

Identifies viable replacement options available today.

Community evidence

The sky is falling. The best value DeepSeek plan is also gone. Note: Moderate programmer usage requires about 100 million tokens per workday. Currently, only mimo v2.5 can meet monthly usage needs. Everything else requires purchasing multiple copies. Since everyone's working hours actually fall within DeepSeek's peak time, the actual monthly token quantities now are: v4pro: 200 million (approximately lasts 2 workdays), v4flash: 680 million (approximately lasts 7 workdays).

04

Claude 20× Pro Burns Through Weekly Limits in Hours; Opus 5 Communication Degrades with Jargon

What Happened

Users on the 20× Pro plan report that a single 23-minute debugging prompt session consumed 99 %–20 % of their weekly allowance, meaning $200 per week lasts only about three working days.

One user noted that $200 per week previously meant essentially unlimited use, but Codex now imposes limits unless running 100 sub-agents.

Claude's authentication service went down; OAuth sessions were terminated and users could not reconnect.

Opus 5 and Fable 5 exhibit communication degradation: they use odd terminology such as 'chips' to refer to UI elements, strange abbreviations like 'server repoint' instead of 'server repointing to the new server,' and persistently list unticked items they claim not to have touched, creating an infinite loop.

Users attempted CLAUDE.md custom instructions for two weeks without fixing Opus 5's 'LinkedIn broetry' writing style.

Some users switched from Claude to Codex and found reliability and output quality on par with older Claude versions.

Community Reaction

Users report that $20 buys up to 50 M tokens per 5 hours on Claude including cache.

Claude 5 series is described as 'lobotomized' — users consider rolling back to Opus 4.6.

Multiple users simultaneously say they are switching back; Sol is good but Opus 5 gets more actual work done.

Users observe that model performance is strong at launch then gets 'detuned,' cutting capability in half within weeks.

Users complain that Claude's safety overreach diminishes practical utility.

Practical Takeaway

The 20× plan's cost-performance is now seriously in question — reassess whether it justifies the spend.

Heavy users may need to consider Codex as a replacement.

Opus 4.6 may be the last reliable Anthropic model.

CLAUDE.md custom instructions have limited effect on improving Opus 5's communication style.

Use Case

Complex projects with heavy token consumption.

Tasks with specific writing tone and communication requirements.

Migration decisions from Claude to Codex.

Practical Value

Helps users evaluate whether to maintain their 20× subscription.

Analyzes Codex as a Claude replacement.

Documents communication degradation patterns in the Claude 5 series.

Community evidence

Past 24h limits disappear with barely any results.

02

Model experience tracking

05

GPT-5.6 Sol Praised as OpenAI's Best Vision Model, Outreasoning Claude-Generated UIs

What Happened

GPT-5.6 Sol is described as the best vision model OpenAI has ever released.

When tested on a visual task, GPT-5.6 Sol's trace revealed algorithmic thinking — identifying start and end points and tracking along the minimal-angle path.

Analysis compared GPT-5.6 Sol's output against hallmarks of Claude-produced UIs: animation serving as affordance versus decoration, and system fonts selected for utility.

The analysis flagged clichés in Claude's frontend-design SKILL.md: warm cream backgrounds with serif display, near-black paired with acid tones.

Users point out that VLMs are designed to fail without tool assistance; requesting 'no Python or tools' essentially mandates inefficiency.

GPT-5.6 Sol correctly solves a line-following puzzle when allowed to assist with code.

Community Reaction

The community has begun noticing recurring cliché patterns in AI-generated UIs.

Users argue that VLM visual evaluation tasks are poorly designed, expecting human-like external computation.

Claude's SKILL.md is seen as a product of out-of-touch design leads.

The distinction between animation as affordance (functional signal) versus decoration is actively debated.

Practical Takeaway

GPT-5.6 Sol may outperform Claude on vision tasks.

Vision evaluations should allow models to use tools — otherwise results are meaningless.

AI-generated UI clichés are now being systematically identified by the community.

Tool assistance improves accuracy on visual reasoning tasks.

Use Case

Complex visual reasoning tasks.

UI analysis and design review.

Spatial reasoning tasks such as line-following puzzles.

Practical Value

Guides model selection for vision-heavy tasks.

Identifies common clichés in AI-generated UIs.

Outlines a sensible strategy for using vision models with tool support.

Community evidence

The second answer is far more revealing than the first: OP: > do you think you did a good job there ChatGPT: > I spent 15 minutes, emitted several fake-sounding “tracing the puzzle” progress updates, and then gave a confident permutation without showing that I had actually followed the lines correctly.

03

Ecosystem and open models

03

Qwen 3.8 27B Matches DeepSeek V4 and GPT Luna on Benchmarks, Runs Locally on 24 GB GPUs

What Happened

Artificial Analysis benchmarks show Qwen 3.8 27B performing neck-and-neck with DeepSeek V4 and GPT-5.6 Luna Max.

Users report that Qwen 3.8 27B at full precision solves medium-difficulty requirements and Leetcode Hard problems with ease, delivering Opus 4.6-level performance.

A gifted programmer — a Top 2 university professor's student — said it cannot match Opus 4.6, calling it 'the moment of equality for all.'

Local deployment benchmarks: RTX 3090 with Q5_K_M GGUF achieves approximately 30–32 tok/s; Strix Halo plus 170HX in a 64 GB USB4 dock reaches 3,000–4,000 tok/s for prefill and 40–60 tok/s for text generation with MTP; 7900 XTX plus MCP delivers 70+ tok/s for text and 31–32 tok/s for complex code.

Qwen 3.8 27B's default overthinking problem was flagged by Simon Willison, with reasoning budget controls emerging as a community mitigation.

The model generates text very fast but in strict coding tasks the MTP mechanism rejects large numbers of draft tokens, reducing effective speed.

A 16 GB 4070 Ti Super can run the INT4 quantized version with a 100k context at 35 tok/s.

Community Reaction

The community is stunned that Qwen 3.8 27B can run an Opus 4.6-level model locally.

Users point out that overthinking is a universal affliction of Chinese models — GLM 5.3 and DeepSeek V4 both consume large reasoning token budgets.

Some argue that abundant reasoning tokens are necessary, not overthinking.

The primary bottleneck is hardware: supporting a 1 million context at 150 tok/s exceeds current consumer capabilities.

DeepSeek V4 Flash reduces code hallucination by upweighting citation sections, but local CPU offloading remains too slow.

Practical Takeaway

Qwen 3.8 27B is the first model to deliver Opus 4.6-level capability on a consumer GPU with 24 GB or more.

Reasoning budget is tunable via --chat-template-kwargs '{"reasoning_effort":"medium"}'.

For text tasks, prioritize max speed; for complex coding tasks, factor in MTP-driven draft token rejections.

INT4 quantization enables 16 GB cards to run the model, albeit at reduced speed.

Use Case

Local code completion and medium-difficulty algorithm problems.

Users who need Opus 4.6-level reasoning but lack GPU budgets.

Balancing speed and quality via reasoning budget tuning.

Practical Value

Assists users in evaluating local deployment feasibility.

Provides real-world speed benchmarks across hardware configurations.

Guides reasoning effort parameter tuning.

Prompt

--chat-template-kwargs '{"reasoning_effort":"medium"}'

Prompt Analysis

Setting reasoning_effort to medium reduces overthinking while preserving sufficient reasoning quality.

For simple tasks like UI tweaks, high reasoning effort wastes time and resources.

Community evidence

Qwen3.8-27B shocked me more than Fable5 did. I ran it at full precision on my self-assembled GPU server, solved some medium-difficulty requirements and Leetcode hard problems with ease, and it indeed performs at the opus4.6 level. Thinking back to mid-March to July (before Anthropic's major ban), I was following various tutorials to top up Claude, insisting on using the opus series models. Honestly speaking, opus4.6 really was impressive back then, genuinely helping me complete a project, and later using opus4.7, gpt5.4, gpt5.5, also successfully passing the project acceptance.