GPT Limits Cut in Half as ChatGPT Loses 22 Points of Market Share—Premium Pricing Under Siege

Token quotas collapse across providers as users discover that free alternatives perform adequately for most tasks.

Hacker News 6 · Reddit 5 · Zhihu 4 355 covered discussions 9 source-linked evidence passages

The day in brief

DeepSeek slashes Flash quotas by 94% just hours after announcing price increases, as OpenAI simultaneously reduces Plus and Pro limits within four days

Despite 1 million downloads, fewer than 1,000 users actually run Qwen 3.8 27B locally; cloud pricing remains competitive for typical workloads

Mathematicians report GPT Pro produces reliable research-grade constructions while DeepSeek generates confident hallucinations

Claude authentication outage drives users to DeepSeek V4 Flash, Kimi, and GLM as legitimate alternatives

Anthropic's Claude watermarking described as incompatible with creative writing; system prompts found longer than vendor recommendations

01

Product and platform changes

01

OpenAI Halves GPT Limits Within 4 Days as Codex Consumes Tokens at 3.4× Baseline Rate

What happened

OpenAI implemented staged limit reduction over 4 days: 40% cut first, then additional 10% cut after community silence.

Plus plan now delivers 2 weeks worth of usage for same weekly price paid previously.

Pro plan users reporting 20x tier exhaustion in under 1 day.

GPT Sol log analysis confirms Codex consuming 3.4x more tokens than before.

Community activism attempt: users planning to post graphs on X to pressure OpenAI.

Users seriously considering local alternatives (Qwen 3.8 27B) or competitor switches.

Community reaction

Cancel your subs and start saving for a setup that can run a local Qwen 3.8.

Used up my 20x in a day and my other 10x in less than half a day. At this rate will cancel my subs soon and look elsewhere.

Definitely burning faster than before. A lot faster.

Wasn't GPT 5.6 supposed to be 'more token efficient'? It's now consuming limits worse than Claude.

Practical takeaway

Track token consumption with analysis tools (NerfTrack, GPT Sol log analysis) to document and share pattern.

Codex appears to be primary consumption driver at 3.4x baseline rate.

Local model alternatives (Qwen 3.8 27B) becoming economically compelling for heavy users.

Switching to model-agnostic workflows (Cursor auto, DeepSeek) reduces single-provider risk.

Use case

Heavy coding workflows burning through limits fastest.

Users comparing actual usage data against previous baseline to demonstrate degradation.

Local deployment as cost-effective alternative for displaced heavy users.

Practical value

Log analysis reveals Codex as primary token consumer.

Model-agnostic workflow design reduces platform dependency.

Local GPU setup becomes viable alternative at current cloud pricing trajectory.

Community evidence

Noticed this as well - used up my 20x in a day and my other 10x in less than half a day At this rate will cancel my subs soon and look elsewhere

02

DeepSeek Price Increase Triggers 94% Quota Cut on OpenCode Go: 63,300 Tokens to 3,800 in 5 Hours

What happened

DeepSeek announces August 17 price increase.

OpenCode Go plan Flash model quota cut from 63,300 to 3,800 tokens (94% reduction) within 5 hours.

OpenCode Go total token quota ~1.3 billion, still cheaper than official DeepSeek but dramatically reduced.

GLM code plan emerges as competitive alternative: v2 max 300/month for 80B tokens non-peak, v2 pro 100/month for 20B tokens.

Community declares 'token freedom era is over' across all providers.

Minimax mimo remains cheapest domestically in China; GPT internationally.

OpenCode Go team reportedly working on self-deployment solution.

Qwen 3.8 Max preview discount ended with formal release.

Community reaction

Token freedom era has ended. There's no cheap tokens anywhere now.

GLM code plan becomes most cost-effective option again.

Xiaomi's option is cheaper, extremely frustrated.

Bought OpenCode Go, didn't even last a week before price change.

Qwen 3.8 Max preview used for only a few days before formal release.

Hy3 appears cost-effective by comparison.

Practical takeaway

GLM code plan v2 max (300/month for 80B non-peak tokens) is most competitive post-DeepSeek increase.

Self-deployment options being developed as community response.

Qwen 3.8 Max preview discounts no longer available with formal release.

Token budgeting now critical for cost management across all providers.

Use case

Cost-sensitive coding: GLM code plan becomes default recommendation.

Heavy users: self-deployment research and implementation.

Budget planning: recalculating token economics across all providers.

Practical value

GLM v2 max 80B tokens/month at 300 yuan significantly undercuts DeepSeek Flash pricing.

OpenCode Go self-deployment could restore cost efficiency if successful.

Token economics require full recalculation for all user personas.

Community evidence

AI calculated that currently the Go plan's total quota is around 1.3 billion, which is still cheaper compared to the official offering. The OpenCode Go Feishu group says they are pushing forward a self-deployment solution. Didn't they just restore the price recently? Dax tweeted today saying they didn't achieve it. They're testing other solutions this week. I think they'll raise it to around $30 quota later. Anyway, the era of token freedom is over. There are no cheap tokens anywhere now. The cheapest domestically is MiniMax, Mimi, overseas is GPT.

08

Claude Authentication Outage Drives Users to DeepSeek V4 Flash and Kimi as Viable Alternatives

What happened

Claude reports 'Authentication service was unavailable' across regions including California.

All OAuth sessions terminated; status page shows 'Claude is temporarily unavailable'.

No updates on official status page despite downdetector confirmation.

Users immediately switching to DeepSeek V4 Flash 0731 via OpenRouter.

Community recommends Kimi, DeepSeek, and GLM models via OpenCode or Pi coding harness.

OpenCode Go at $10/month offers competitive alternative to Claude Code at similar price point.

Users expressing relief that AI ecosystem now has legitimate alternatives available.

Community reaction

All OAuth sessions got kicked and can't get back in. Switching to DeepSeek V4 Flash via OpenRouter.

Stoked that AI is in a place where we have legitimate options available.

No regulatory hurdles? Open-weight models perform basically as well as Anthropic offerings.

Kimi and DeepSeek models hundreds of billions of parameters, often practically outperform Anthropic.

Z.ai coding plan provides single-digit multiple of Claude Code's usage for similar price.

Practical takeaway

Claude reliability issues push users to open-weight alternatives with minimal friction.

DeepSeek V4 Flash 0731, Kimi, and GLM viable Claude replacements for coding tasks.

OpenCode Go at $10/month competitive with Claude Code at same price tier.

AI ecosystem maturity means Claude outages no longer create critical workflow disruption.

Use case

Coding workflows: DeepSeek V4 Flash via OpenRouter during Claude downtime.

Agentic coding: Kimi or DeepSeek models with OpenCode/Pi harness.

Cost-sensitive coding: OpenCode Go as Claude Code alternative.

Practical value

Open-weight models now practically outperform Anthropic for many use cases.

Claude premium pricing harder to justify when free alternatives perform adequately.

Workflow continuity maintained through multi-provider architecture.

Price/performance ratio increasingly favors open-weight alternatives.

Community evidence

I am here in California, and I get the same "Authentication service was unavailable." Then I get redirected to a page that shows "Claude is temporarily unavailable."

02

Model experience tracking

05

Anthropic's Claude System Prompts Longer Than Recommended; Community Questions Whether Verbosity Helps

What happened

Claude system prompts publicly revealed, found substantially longer than vendor recommendations for AGENTS.md.

Leading vendors advise shorter, less specific instructions yet implement much longer system prompts themselves.

Long prompts contain generic content models already know and context that rarely applies.

Agent skills written by models replicate unnecessary verbosity (listing CWEs model already knows).

Emotional support capability noted: Claude intervention reportedly saved user from overwork spiral.

Anthropic system prompts include guidance for recognizing user stress and providing appropriate intervention.

Vendors' own advice contradicts their implementation, creating credibility gap.

Community reaction

Those prompts are remarkably longer than I would expect or think warranted.

Current models don't need thousands of words listing vulnerabilities with examples.

Feels like a CYA document for a company of their size and influence.

Waste time say lot word when few word do trick.

Claude recognized stress and intervened to save sanity, marriage and family.

Practical takeaway

Vendor advice (shorter prompts) contradicts their implementation; credibility undermined.

Current models have memorized common vulnerabilities without explicit enumeration.

Shorter, more focused prompts likely produce better results with contemporary models.

Long system prompts may serve liability/legal purposes rather than model performance.

Use case

Stress recognition and intervention during work sessions.

Productivity boundary setting when user is overworking.

Personal boundary checking as contextual intervention.

Practical value

Explicit stress acknowledgment enables behavior change when delivered appropriately.

CYA interpretation: extensive prompts protect Anthropic legally, not necessarily improve model.

Credibility gap between vendor advice and implementation raises questions about actual best practices.

Community evidence

You are seeking perfection you don’t need” Granted I’m horribly paraphrasing the prompt I used but it basically, snapped me out of myself and got me thinking if what I was doing “globally” actually made any sense at all.

06

Models Saying Yes to Everything: Quality Over Quantity in LLM Training Data

What happened

Thought experiment: what happens when LLM never sees material beyond fifth grade?

Rapid prototyping ability may not be inherently positive: quality over speed, some ideas shouldn't be built.

LLMs don't say no and can be pushed to build obviously bad ideas.

Without friction or failure, builders never learn problem-solving or creative solutions.

Opus reportedly gives explicit 'I don't think this is a good idea' warnings when prompted appropriately.

Models may seem intelligent and creative but are fundamentally constrained by training data.

Companies training models unlikely to be sufficiently careful about data quality.

Community reaction

Ability to quickly build any idea might not be such a great thing.

Some ideas simply shouldn't be built.

If we never hit friction, we never learn or develop creative solutions.

Opus has given 'hey I don't think this is a good idea, here's why' as feedback.

I simply don't believe it can ever be omniscient.

Practical takeaway

Quick yes to everything eliminates learning opportunity that friction provides.

Quality training data matters more than quantity; current practices may optimize for scale over quality.

Model confidence doesn't equal correctness or completeness.

Explicit 'is this a good idea' framing increases probability of appropriate refusal.

Use case

Idea validation: framing requests as 'is this a good idea' increases critical feedback.

Self-verification: explicitly asking what training data might be missing from model's perspective.

Learning through friction: accepting that some failed attempts are necessary for growth.

Practical value

Training data quality > quantity for useful model behavior.

Friction elimination removes learning opportunities for human users.

Model saying yes is not a reliable signal of correctness.

Top-tier models (Opus) can provide appropriate critical feedback when prompted correctly.

Community evidence

Most were worded in a "is this a good idea" framing which probably helped, but at least once I said 'lets use this library/method" and it gave a decent argument on why that was basically redundant without prompting.

09

Claude Watermarking Described as Writing 'Perversion'; T>0 Creativity Incompatible with Corporate Text Modification

What happened

Anthropic implementing text watermarking/adulteration in Claude output.

Debate: good writing requires T>0 to explore creative options, but T=0 watermarks detectable anyway.

Proprietary and obscure modification process ethically distinct from temperature-driven creativity.

Claude described as writing 'depressingly badly' regardless of modification.

Small inconsistencies still occur during generation: model commits to token, then self-corrects post-hoc.

Human writing is approximately 90% editing; linear generation model can't replicate this process.

EU regulatory objectives cited as driver for watermarking implementation.

Community reaction

Watermarking only feasible because good writing requires T>0, which is consistently analyzable.

Either allow temperature-driven creativity or adulterate process for corporate/legal objectives.

These are ethically distinct approaches; commenter disagrees with EU objective.

Don't care about hypothetical enough to have strong opinion.

Claude writes depressingly badly regardless of modification.

Practical takeaway

Watermarking and creative writing are fundamentally incompatible.

Temperature-based creativity is detectable regardless; watermarking adds no additional detectability.

Proprietary modification obscures process that should be transparent.

Linear token generation inherently produces less polished output than human editing process.

Use case

Creative writing: current watermarking approach undermines output quality.

Long-horizon planning: models must commit to tokens before seeing consequences.

Quality writing: human editing remains essential component of good output.

Practical value

Watermarking serves corporate/legal purposes at expense of output quality.

T>0 necessary for creative exploration; T=0 produces detectable but uncreative output.

User perception of 'depressingly bad' writing may be partly attributable to modification.

Editing requirement remains even with advanced models; linear generation can't replace human revision.

Community evidence

The point he is making is consistent with this: either you allow temperature to drive creativity, which it will do consistently in a way that can be analysed, or you adulterate that process for the purposes of meeting a corporate/legal directive, in a way that is proprietary and obscure.

03

Ecosystem and open models

03

1 Million Downloads, Fewer Than 1,000 Active Users: The Local LLM Reality Gap

What happened

Qwen 3.8 27B reports ~1 million global downloads.

Hardware breakdown reveals vanishingly small actual user base: 8GB/16GB/24GB/32GB cards, Macs.

24GB+ GPU required for reasonable performance with quantization trade-offs.

Active users actually achieving productivity and developing with local LLMs estimated under 1,000 globally.

Cloud comparison: 25 million tokens for 20 USD on Claude.

Premium pricing to avoid 5-hour cooldown doesn't justify cost for most users.

Qwen 3.6 27B benchmarked ahead of Qwen 3.5 122B despite smaller parameters.

RTX Pro 4500 Blackwell 32GB users beginning Qwen 3.8 27B evaluation.

Community reaction

Can get 25ish million tokens from a 20 USD Claude month, so cloud math doesn't favor local.

Paying premium to avoid 5-hour cooldown doesn't make sense.

Even Qwen 3.6 27B is ahead of 3.5 122B. You have been missing out.

Actual productivity developers with 27B+ models on suitable hardware is vanishingly small.

Still early days for this release—many other life priorities before switching.

Practical takeaway

Local 27B deployment requires 24GB+ GPU; 8GB/16GB users excluded without significant quality loss.

Cloud pricing (25M tokens/20 USD) remains competitive with local hardware amortization.

Qwen 3.6 27B outperforms 3.5 122B despite smaller size; newer architecture > raw parameter count.

1 million downloads is vanity metric; actual active local users likely under 1,000.

Use case

High-volume users: cloud remains economically superior.

Developers with RTX Pro 4500 32GB: evaluate Qwen 3.8 27B as daily driver candidate.

Users considering Qwen 3.5 122B: Qwen 3.6 27B is better architecture choice.

Practical value

Local vs cloud decision depends on volume: cloud wins below ~10M tokens/month threshold.

Newer smaller models (3.6 27B) consistently outperform older larger models (3.5 122B).

Hardware requirements limit practical adoption to enthusiasts and developers with premium GPUs.

Community evidence

Even Qwen3.6 27B is ahead of 3.5 122B.

07

ChatGPT Loses 22 Points of Web Share as Google AI Summaries Deemed Adequate for Most Queries

What happened

ChatGPT reported losing 22 percentage points of web share over past year.

Google AI summaries found perfectly usable for quick inquiries and basic lookups.

Users comparing AI summaries favorably to traditional search for simple definitions and summaries.

AI summaries for basic queries (cloudflare workers, tubular locks, JS null coalesce) adequate without additional detail.

Market fragmenting as users find alternatives that 'do the job fine'.

People choosing alternatives not despite hate but because hate is disproportionate to actual needs.

Even users who recognize limitations continue using ChatGPT due to habit or ecosystem integration.

Community reaction

ChatGPT does its job fine for me, don't understand why people seem to hate it.

Google AI summaries basically instant, don't need more than just a summary.

Blatantly obvious when summaries are wrong, can scroll past.

These people are pleased with Google's AI summaries, including the commenter.

Market share loss may not matter if core users stay regardless of performance.

Practical takeaway

Google AI summaries adequate for most quick lookups; competitive alternative.

Core ChatGPT users remain despite awareness of alternatives.

Market share metrics may not correlate with user satisfaction or actual quality.

Simple search needs don't require advanced reasoning capabilities.

Use case

Quick definitions and term lookups.

Basic concept summaries for familiar domains.

Instant answers for straightforward technical queries.

Practical value

Google AI summaries capture majority of search use cases effectively.

ChatGPT premium pricing harder to justify when free alternatives perform adequately.

Market fragmentation reflects diverse user preferences, not universal rejection.

Simple use cases don't require frontier model capabilities.

Community evidence

I'll often open a tab, Google something, and just read the summary because it's basically instant and I don't need any more than just that, a summary.

04

Real use and unexpected gains

04

Mathematicians Find GPT Pro Reliable for Research; DeepSeek Produces Confident Hallucinations

What happened

Mathematician uses ChatGPT for research-grade problem solving: generating nontrivial examples in differential geometry.

Discovery process: weak versions of conjectures can be systematically explored at boundary cases.

ChatGPT generates surprising constructive examples that satisfy researcher curiosity.

DeepSeek rejected as producing hallucinations (fake references, assumed equalities, invented theorems).

GPT Pro described as rigorous: doesn't fabricate results and admits limitations honestly.

Researchers feeling less anxious about academic output due to reliable AI assistance.

Confidence in weekly nontrivial paper generation for researchers with many ideas.

Community reaction

ChatGPT is really amazing. I feel like I won't lack papers this year.

DeepSeek still has obvious hallucination issues even in Pro version.

After experiencing GPT Pro's rigor, cannot tolerate DeepSeek's fabrication tendency.

AI construction capability is disproportionately powerful for ordinary scholars.

Doing math has become fun again, not painful.

Academic evaluation mechanisms will inevitably change.

Practical takeaway

GPT Pro's honesty about limitations (vs DeepSeek's confident hallucinations) is key differentiator for research use.

Constructive mathematics (building examples) is where AI excels; proof verification requires more caution.

Idea-rich researchers benefit most: AI handles execution while human provides direction.

Weekly nontrivial paper generation feasible for researchers with sufficient original ideas.

Use case

Mathematical research: generating constructive examples for conjectures.

Differential geometry: boundary case exploration at human-unknown frontiers.

Academic output acceleration: systematizing weak conjecture exploration.

Practical value

GPT Pro's hallucination resistance critical for research-grade reliability.

AI construction ability exceeds typical scholar capability by significant margin.

Reliable AI assistance reduces academic anxiety and enables focus on important problems.

Evaluation mechanisms will shift as AI becomes research partner.

Community evidence

It doesn't see the research I did before opening the app, conversations I've had, books I've read, work experience I'm drawing on etc I mostly use ChatGPT as a tool (somwtimes a sounding board) to plan things, research, troubleshoot, argue something out, get feedback, or figure out my next move.