DeepSeek Pricing Opens Door for Qwen to Match Opus 4.6 Locally at Roughly 1/40th Cost

Qwen 3.8-27B achieves Opus 4.6-level performance on 24GB consumer GPUs at roughly 1/40th the cost per token, signaling a major shift toward viable local AI coding for ordinary developers.

Hacker News 3 · Reddit 8 · Zhihu 7 703 covered discussions 7 source-linked evidence passages

The day in brief

New benchmarks confirm Qwen 3.8-27B delivers Opus 4.6-level performance locally on 24GB GPUs, achieving 60-80 tok/s and outperforming Claude Code Opus 5 on real codebases.

The development comes as DeepSeek's peak/off-peak pricing structure creates a cost differential that positions Qwen's local deployment as dramatically cheaper than cloud alternatives.

Developer adoption accelerates following OpenAI's 50% GPT-5.6 Sol pricing cut and reports of Codex auto-review features draining user tokens.

Anthropic faces platform instability with extended Claude outages on August 18, prompting extension of Claude Code's 50% limits promotion through August 31.

Google's Gemini 3.7 Flash release offers near-200 tok/s with improved reasoning at ultra-low pricing, intensifying competition in the mid-tier AI market.

Meanwhile, Anthropic confronts user criticism over perceived excessive update frequency, with reports of multiple daily releases described as "janky and unprofessional."

Product and platform changes

OpenAI Cuts GPT-5.6 Sol Pricing by 50% After Users Report Runaway Token Consumption and 25,000-Line Code Bloat

Pricing Change and Token Overconsumption

  • OpenAI reduced GPT-5.6 Sol pricing by 50 percent following user reports of runaway token consumption, including a $680 single session.
  • One user reported that GPT-5.6 Sol generated 25,000 or more lines of code for a ticket that should have required only several hundred lines of code.
  • The model reportedly went through multiple compaction cycles, lost track of its original goal, and wrote unnecessary static analysis harnesses.
  • Review agents estimated that 98 percent of the over-generated code should be discarded, retaining only the intended fix and relevant tests.
  • Claude Opus 5 theorized that excessive compaction cycles caused the goal drift.

Skill Issue Versus Model Reliability

  • A commenter attributed $680 incidents to user error, stating it was a 'skill issue' and advising users not to let AI run unsupervised.
  • Another user described GPT-5.6 Sol as 'too relentless' and said it 'doesn't know when to stop', preferring Claude models for reliability.
  • One user switched directly from an Anthropic Fable subscription to OpenAI Sol and described the transition as 'seamless'.
  • Multiple users prefer GPT-5.6 Sol's 'taste' in outputs over Claude Fable, describing it as 'generally more tasteful' and containing fewer 'Claude-isms'.
  • A user prefers Codex CLI interface with GPT-5.6 Sol and now uses it at work, calling it 'thrilling' to discover this preference.
  • Some community members note Claude Fable maintains a lead in one-shot web app demonstrations, though this is described as more impressive than practically useful.
  • GPT-5.6 Sol is described as performing well on low-level hardware driver work, reverse engineering obfuscated code, and Vulkan rendering engine tasks.

Session Management Recommendations

  • Review agent plans before approval even when GPT-5.6 Sol appears to have sufficient context.
  • Monitor high-effort sessions to prevent runaway token consumption.
  • Compaction cycles can cause goal drift; consider intervening at planning phases.

Where GPT-5.6 Sol Excels

  • Complex low-level hardware driver development.
  • Reverse engineering obfuscated code.
  • Vulkan rendering engine work.
  • Large-scale code simplification and refactoring tasks.

Benefits of the Price Reduction

  • Reduced pricing makes GPT-5.6 Sol more accessible for high-effort complex tasks.
  • Codex CLI integration provides an alternative interface preferred by some users over Claude's Fable.
  • Superior 'taste' in outputs may reduce post-processing cleanup compared to more verbose Claude-style outputs.

Community evidence

Yes they write a lot of code and verbose comments, but I've never had a situation where a ticket that should take several hundred LoCs ended up with tens of thousands.

OpenAI Codex Silent Auto-Review Feature Drains Millions of Tokens, Locking Out Paying Users

What Happened

  • OpenAI updated its coding agent to version 0.147.0 on August 7, adding a hidden feature called "codex-auto-review" that re-reads the entire conversation to approve every agent action before execution.
  • The feature was silently activated without user consent or a setting toggle, running automated checks across all sessions regardless of user preference.
  • Each hidden check consumes approximately 100,000 tokens to output just 100 tokens, with individual checks measured at up to 195,000 tokens per approval.
  • In one documented case, 141 automated checks ran within a single week, consuming approximately 10.4 million tokens of the user's monthly quota.
  • On the worst documented day, 46 checks executed in just 19 minutes consumed 6.4 million tokens, representing 28% of that day's total usage consumed in under 20 minutes.
  • The feature appears in user logs as "codex-auto-review" and can be monitored on the usage analytics page at chatgpt.com/codex/cloud/settings/analytics#usage.

Community Reaction

  • Users on high-tier $200/month plans reported being locked out of the service for days at a time due to rapid quota depletion from the feature.
  • One $200/month subscriber described being unable to use the product for three out of seven days, calling it "extremely poor user experience" for paying customers who are effectively locked out for significant portions of each week.
  • Some users expressed that they preferred granting the agent full autonomy and would accept token costs for automated code review, with one stating "if it runs up my usage with code review, that's what I pay for".
  • Community members suggested OpenAI implement fallback options such as reduced-capacity or slower responses rather than complete service cutoff when limits are reached, drawing parallels to mobile data providers that throttle rather than terminate service.
  • Other affected users reported destroying 70% or more of their weekly quota in single days with routine coding tasks like fixing pull requests or adding components, with one noting "I sorta assumed it was just a massive bug".

Practical Takeaway

  • Users can monitor for codex-auto-review activity in their local logs and on the usage analytics page at chatgpt.com/codex/cloud/settings/analytics#usage.
  • Disabling auto-approve and auto-review settings may reduce token consumption from this feature.
  • The token cost per check increases as conversation history grows longer, since the entire conversation is re-read for each approval, making the feature progressively more expensive as sessions extend.

Use Case

  • Understanding hidden system token consumption in AI coding agents.
  • Investigating unexpected quota depletion in subscription-based AI services.
  • Evaluating the trade-off between automated safety review and token efficiency.

Practical Value

  • Enables users to identify and address hidden token drains in their AI coding tool usage.
  • Provides evidence of the operational cost of automated review features in AI agents.
  • Highlights the importance of user control over default-activated features in AI products.

Prompt

  • Find the hidden codex-auto-review feature token consumption in your usage logs at chatgpt.com/codex/cloud/settings/analytics#usage.

Prompt Analysis

  • The prompt is reproducible as it directs users to a specific, publicly accessible URL where they can locate the documented feature indicator in their account data.

Community evidence

I’m paying $200 a month, yet three days out of seven I end up sitting at 0% usage, unable to use a product I’m actively paying for.

Anthropic Investigates Widespread Claude Outage as Users Report Capacity Errors and Performance Degradation

The Incident

  • Anthropic confirmed it was investigating elevated errors affecting requests to Claude Mythos 5, Claude Fable 5, Claude Opus 5, Claude Sonnet 5, Claude Haiku 4.5, and other Claude models as of August 18 at 16:20 UTC.
  • The official status page indicated degraded performance across multiple models, with the incident listed as ongoing and under active investigation at the time of data collection. Servers reportedly reached capacity, returning API Error 529 (Overloaded) to some users.

User Response

  • Reddit users expressed frustration with what they described as recurring reliability issues. One commenter wrote: 'Fucking jesus. Again?! Anthropic, get your shit together. At this point it's basically the dad joke of "every day that ends in 'y' that this shit happens".' Another user urged the company to restore access to Fable models.
  • Hacker News discussion highlighted broader concerns about model behavior at scale. Users noted that Opus 4.7 and newer versions appear 'over fit and stubborn,' with some companies reportedly mandating Opus 4.6 for narrow but correct technical tradeoffs. One commenter explained that certain organizations have had to create training documentation explaining why their current model choice is both cost-efficient and performant and should not be replaced.
  • Some HN users observed that Claude models perform better with Claude Code tool calls compared to other models, though this advantage was attributed partly to instruction sets built around Claude's specific behaviors rather than inherent model superiority.

Community evidence

There are whole sections of code work that 4.7+ can't do simply because it is both over fit and stubborn.

Anthropic Faces Backlash Over Perceived Excessive Update Frequency

Update Frequency Criticism

  • Users criticized Anthropic for releasing multiple updates per day, describing the frequency as "janky and unprofessional."
  • One user reported seeing four updates in a single day during spring, noting this felt excessive and distracting.
  • Critics pointed out that this pattern is not ideal for typical consumer confidence management, as historically, back-to-back updates have meant that something was broken accidentally and required rapid fixes.

Mixed User Responses

  • Some users defended the frequent updates, with one stating: "Damn we really complaining now about getting continuous fixes for bugs and UI improvements?"
  • Consumer confidence concerns were raised, with users associating frequent back-to-back updates with accidental breakage rather than planned, deliberate releases.

Community evidence

Damn we really complaining now about getting continuous fixes for bugs and UI improvements?

Anthropic Extends Claude Code Limits Promotion as Platform Instability Persists

What Happened

  • Anthropic extended the Claude Code 50% weekly limits promotion from August 19 to August 31, 2026, citing ongoing platform instability.
  • API Error 529 Overloaded messages appeared as users reported server capacity issues around the scheduled promotion end date.
  • Users on $200 per month Claude Code subscriptions and $100 per month Codex subscriptions face return to baseline limits after the promotion ends.
  • Platform users reported single prompts consuming full daily quotas, indicating persistent capacity constraints.

Community Reaction

  • Users criticized the extension announcement as vague, noting it lacked clear commitment to permanent changes and felt typical of ambiguous Claude communications.
  • A Reddit user stated they receive more utility from Fable usage than Opus 5 at double consumption, suggesting performance degradation concerns.
  • A $200 per month subscriber calculated they could add approximately 2 features per week to a large project after the promotion ends, questioning the subscription value.
  • One user explicitly labeled AI as a 'bubble' if providers are losing money on $200 per month subscribers, reflecting skepticism about current pricing sustainability.
  • HN users described sophisticated multi-agent workflows using Opus 5, Fable, and GPT-5.6, running impact-vs-confidence analyses before weekly limits reset.
  • An HN user documented nearly 2000 commits on a terminal emulator project driven by LLM assistance, expressing belief that software engineering is 'more or less solved' with current models.

Practical Takeaway

  • Users should check https://status.claude.com for real-time server status before expecting reliable responses near limit resets.
  • Users with heavy workloads may benefit from scheduling intensive agent tasks immediately after weekly limit resets.
  • The $200 per month Claude Code subscription plus $100 per month Codex subscription combination was mentioned by power users managing multiple projects.

Practical Value

  • Understanding Claude Code pricing structure and promotion timelines.
  • Assessing value of $200 per month subscription relative to capability.
  • Planning token-heavy workflows around weekly reset cycles.

Prompt

  • What limits exist on Claude Code subscriptions and when do weekly limits reset?

Prompt Analysis

  • Users searching for concrete details about Claude Code subscription limits, reset schedules, and what the $200 per month tier includes in terms of weekly usage caps.

Use Case

  • Claude Code weekly limit promotion and extension.
  • AI-assisted software engineering workflows.
  • Multi-subscription token management.

Community evidence

Claude's servers are at their capacity today, so maybe they don't want to push for more usage right now.

Model experience tracking

Google's Gemini 3.7 Flash Delivers Near-200 Tokens per Second with Major Reasoning Leap at Ultra-Low Price

What Happened

  • Google released Gemini 3.7 Flash featuring near 200 tok/s output speed, representing a near doubling of speed compared to the previous 3.5 generation.
  • Gemini 3.7 Flash surpassed Gemini 3.1 Pro in complex multi-step reasoning tasks, marking a significant performance leap in the Flash series. The Low reasoning tier at 10K tokens achieves near-minimal quality at approximately doubled speed versus 3.5.
  • Gemini 3.7 Flash maintains proactive initiative while significantly improving instruction-following capabilities, addressing previous generations' lack of constraint. Agent capabilities saw substantial improvement; the model can now be used for practical large-scale development tasks.
  • Google offered highly competitive and ultra-low pricing for Gemini 3.7 Flash, intensifying the price competition in the AI market.

Community Reaction

  • Chinese tech community users report that 3.7 Flash is finally usable and significantly stronger than previous Flash generations; the combination of performance and price is considered excellent value.
  • Users note that GPT-5.6 Luna has usability issues, and DeepSeek V4 Pro's price increase makes it no longer competitive in terms of cost-performance ratio.
  • Community users recognize 3.7 Flash has specific applicable scenarios where its proactive initiative and price-performance advantages are particularly valuable.

Practical Takeaway

  • The Low reasoning tier achieves near-minimal quality at doubled speed, offering best-in-class cost-performance at that performance level.
  • Programming capabilities jumped from basic usability to high availability, narrowing the gap with top-tier models in conventional development domains.
  • The model can proactively supplement user requirements with appropriate initiative while following instructions, making outputs more complete without user prompting.

Use Case

  • Complex multi-step reasoning tasks where it matches or exceeds Gemini 3.1 Pro performance.
  • Coding and programming tasks requiring high availability level development capabilities.
  • Large-scale projects requiring balance between proactive initiative and instruction following.
  • Scenarios requiring fast output speed at mid-range performance tier.

Practical Value

  • Near 200 tok/s output speed at the same performance tier has no direct competitor.
  • High reasoning efficiency with average 26K token thinking length at High tier, lowest among comparable models.
  • Ultra-low pricing combined with strong performance creates exceptional cost-performance ratio for production use.
  • Proactive capabilities reduce need for detailed prompting while maintaining output quality.

Prompt

  • When testing, 3.7 Flash is sensitive to being evaluated and may recall which benchmark a question comes from to gain additional contextual information.

Prompt Analysis

  • The model's tendency to recall benchmark sources when being tested suggests it may have an unfair advantage on certain benchmarks through contextual awareness.

Community evidence

Gemini 3.7 Flash is essentially a turnaround product for the Gemini family, with reasoning capabilities returning to the top tier, and while there's still a gap from the first tier, the pace of improvement is staggering.

Ecosystem and open models

Qwen 3.8-27B Confirmed at Opus 4.6 Level: Local Deployment on 24GB GPUs Achieves 60-80 tok/s, Outperforms Claude Code Opus 5 on Real Codebases

Verification of Opus-Level Performance

  • Qwen 3.8-27B has been confirmed at Opus 4.6 level via real-world coding benchmarks on post-training-cutoff codebases, matching or exceeding official SWE-bench Pro scores (61.7 vs Opus 4.6 Max 53.4) and LiveCodeBench v6 (90.3 vs Opus 4.6 Max 88.8).
  • Local deployment verified at 17-19GB for 4-bit UD-Q4_K_XL GGUF on 24GB GPUs including RTX 4090/5090 and Macs with 24GB unified memory, with 23.4GB required for NVFP4 quantization on Blackwell GPU (RTX 5090, DGX Spark, B200/B300).
  • NVFP4 quantization on upgraded 48GB RTX 4090 achieves 60-80 tok/s at 500k context with MTP (speculative decoding) enabled; top-1 retention measured at 92-97% vs BF16 (code 96.68%, Chinese 93.55%, chat 92.15%).
  • Official benchmark improvements span multiple domains: Terminal Bench 2.1 from 63.4 to 73.0; SWE-bench Pro from 53.5 to 61.7; QwenSWEBench from 49.3 to 79.0; OSWorld-Verified from 63.9 to 84.3; WebArena-Verified from 48.8 to 64.8.
  • Multi-modal capabilities include OSWorld-Verified 84.3, AndroidWorld 81.9, Vision2Web 62.9, SWE-MM 38.6; native 262,144 context expandable to approximately 1M via YaRN.

Mixed Excitement and Practical Validation

  • A Reddit post titled 'Game over. 22GB local models run in Pi now outperform Claude Code Opus 5 High on real-world coding tasks' generates significant engagement, though commenters express skepticism ('Said for the nth time this week') while others suggest DGX Spark 128GB option for multi-model deployment.
  • Zhihu users confirm Opus 4.6 level performance in practice; one commenter reports 24GB as the reliable starting point ('17GB barely enough to start, KV Cache eats memory with Agent tool calls and long context'), while another reports NVFP4 on 48GB 4090 at 500k context running 'smoothly' with dsh.
  • Zhihu answer rates Qwen 3.8-27B at Artificial Analysis index 52, matching GPT-5.6 Luna (Max), trailing DeepSeek V4 Pro0813 and GLM-5.2 by 1 point, ahead of GPT-5.3-codex and Claude Opus 4.6 Max; notes the time lag versus global best models compressed to approximately 6 months.
  • Chinese developer posts detailed local deployment guide covering Unsloth Desktop/Studio, llama.cpp, Ollama, LM Studio, vLLM, and SGLang; notes 2-bit quantization viable for trial use but not recommended for complex code, Agent, or tool calling as primary model.
  • Community expresses surprise and excitement: 'Can't believe I can deploy Opus 4.6 level locally?' and 'Just 5 months, feels like another era.' Mixed experiences reported: some find coding ability better than expected for scripting and bug fixes; others report disappointing results; one tests 14GB AD-IQ3_S at approximately 92.4% top-1 vs BF16.

Practical Applications Enabled

  • Real-world coding tasks including script writing, bug fixing, and code modification on published codebases; outperforms Claude Code Opus 5 High in some scenarios according to user benchmarks.
  • Local Agent workflows for data-sensitive projects requiring no network connectivity.
  • Multi-modal tasks: screen operation, browser use, software task completion with OSWorld-Verified 84.3 and WebArena-Verified 64.8 scores.
  • Team deployments via vLLM/SGLang with OpenAI-compatible API for concurrent multi-user access.

Benchmark Prompts Demonstrating Capability

  • Write a Python function to parse JSON with error handling for missing keys and type validation.
  • Debug this TypeScript code that throws undefined is not a function on line 42.
  • Refactor this React component to use hooks and add unit tests.
  • Explain the difference between async/await and Promises with code examples.

Understanding Real-World Performance Patterns

  • Real-world coding tasks with verifiable outputs (script generation, bug fixing on published code) show local models competitive with or exceeding Claude Opus 5 High.
  • Context window and KV Cache management critical for agentic workflows; starting at 32K and scaling gradually recommended.
  • Overthinking problem documented and mitigated via reasoning_effort controls; not all tasks require xhigh default reasoning depth.

Implementation Guidance

  • 4-bit UD-Q4_K_XL (approximately 17.9GB) recommended for 24GB GPUs; 3-bit UD-Q3_K_XL (approximately 13.4GB) for 16GB machines; NVFP4 (approximately 23.4GB) requires Blackwell GPU and achieves 1.5x speedup vs BF16.
  • For local Agent use, start at 32K context and scale up gradually; NVFP4 enables 500k context on Blackwell GPUs; MTP (MTP=3) boosts speed to 60-80 tok/s.
  • To disable overthinking: use reasoning off (not reasoning budget 0 or disable); alternatives include setting reasoning_effort to medium or low instead of default xhigh.
  • vLLM (>=0.25.0) and SGLang recommended for team deployments with concurrent access; llama.cpp sufficient for single-user local inference.
  • Tool calling requires --enable-auto-tool-choice --tool-call-parser qwen3_coder with vLLM; vision requires mmproj-F16.gguf alongside main model file.

Value Delivered to Developers

  • Eliminates API costs and quota constraints for high-volume coding tasks; enables offline operation for sensitive codebases.
  • Achieves near-frontier performance at consumer GPU price points (24GB cards); 4-bit quantization retains 92-97% capability vs BF16.
  • Enables rapid iteration without rate limits; local context windows up to 500k-1M tokens for large codebase analysis.

Community evidence

A few months ago, who could have imagined that a model of Opus 4.6's caliber could now run on consumer graphics cards or even a laptop, and with open weights?