← All issues

/ COMMUNITY DAILY

What the Community is Discussing Today

Read the hottest discussions from each community: specific experiences, differing viewpoints, and follow-up updates.

4 communities·12 conversations·19 min read

In this issue
01

Reddit

3 selected conversations

Codex Reopens $200 Pro Subscription: Halved Usage Quota Draws Strong Criticism

On September 29, Codex CEO Tibo posted that the $200 Pro subscription would be reopened with adjusted usage calculations. He explicitly stated that the new plan's API spend quota is halved compared to the old Pro $200, i.e., "if you do the math, it will net out at half the dollar in API spend." However, he then listed four reasons trying to prove users would not actually get less: no restoration of the 5-hour limit (ensuring users can fully utilize the weekly quota), subscriptions continuing to yield more work output with higher quality as model efficiency improves, a 50% price reduction for GPT-6 Sol and GPT-6 Luna this week with commitments to continue lowering prices and improving model capabilities, and additional features that do not consume quota to be announced the next day. He said he wanted to be transparent before the big announcement, and more compelling content would follow the next day.

The comments section was almost unanimously critical. Users pointed out that the original post's shift from concrete usage figures to vague promises like "guaranteeing more work completed" is typical corporate rhetoric, using "more value" to mask the fact that the actual quota is halved. One user directly used three AI models (Gemini, ChatGPT, Claude) to summarize the same thing, and all three expressed the same conclusion in different wording: users spending $200 receive half the value of computing resources, while cheaper models and unknown additional benefits are merely used to make this change look like a good thing.

Someone mocked that the "calculation" was nothing more than simple 1 × 50%, questioning why such an obvious halving needed to be packaged with "doing the math." Another reply pointed out that using content generated by the company's own models to expose the company's manipulation tactics was quite ironic. One user said this was "the easiest subscription cancellation ever."

A few users who came forward with personal experience further expressed their dissatisfaction. Someone mentioned that their 20x Pro plan (purchased before it was discontinued) would be exhausted in a single day, and the current model performance was poor with slow responses, directly stating it was "completely useless."

Original thread

Anthropic's Security Research Becomes GLM Ad, LocalLLaMA Discusses Compute Concentration and Advantages of Open-Weight Models

The post title gets straight to the point: Anthropic's blog post about security audits for large language models inadvertently became the most powerful endorsement for GLM (Zhipu AI). The original poster opened with a brief observation—"I always knew GLM was strong, now everyone knows it."—followed by users adding that GLM is not on any government blacklist, requires no special applications, has no nationality restrictions, and won't be banned due to misunderstandings, "no fluff, the model just performs well," with API costs being mere change compared to Claude.

The discussion immediately split into two threads. One continued the "compute concentration" topic: replies attributed the security incidents at Anthropic and OpenAI in recent years to both companies having enough GPUs to launch large-scale attack traffic, while GLM has not appeared on any government vulnerability lists or been hacked, with users even mentioning that GLM was used to defend against closed-source model attacks on Hugging Face. The other thread directly targeted Anthropic: multiple users questioned Dario Amodei's motives—marketing AI penetration testing services and infiltrating enterprises and government agencies publicly while calling for restrictions on open-weight models, essentially a "protection racket" style of competition.

yea bro, I knew GLM was cool.

Skepticism about Anthropic's blog post itself also appeared in the discussion. One user specifically quoted the phrasing "our team has never attempted this type of task before" with sarcastic tone, implying questionable credibility. Multiple replies grouped Anthropic, OpenAI with Google, Meta, and Elon Musk, saying they are "forever on their own blacklist," citing compute concentration risks and their operating styles.

There were also voices pointing out that Anthropic's article overlooked a key fact: the reason enterprises and developers chose GLM for security audits is that Fable lacks deep audit capabilities, implying GLM succeeded through actual performance rather than open policies alone.

Overall, the sentiment of this discussion was clear: most participants viewed Anthropic's security research as reverse publicity, reinforcing preferences for open-weight models while deepening suspicion of closed-source companies' motives, but the post itself did not represent industry consensus or technical conclusions.

Original thread

OpenAI Reopens Codex $200 Subscription Plan, Users Widely Question Shrinking Benefits

OpenAI announced on September 29 the reopening of new user registrations for the Codex $200 subscription plan (link to official announcement tweet). The news sparked predominantly criticism in the r/codex community. Some directly questioned whether this means a major "nerf," pointing out that the previous tier's usage allowance was approximately one-tenth of the API price, which may now be further reduced.

Multiple users mentioned that officials promised to compensate for the usage loss through two-fold model efficiency improvements, but this was generally seen as just empty talk—"that's a bunch of nonsense." Others added that if Astra's costs haven't decreased, users are actually getting less value.

I have altered the deal, pray I don't alter it further

One user described their testing experience in detail: running 5 Opus conversations in parallel, trying to max out a 5x usage plan before the reset date, but only consumed 5% in 5 hours. They said their 20x plan is now "barely moving." In comparison, they mentioned that Anthropic's Claude Max offers higher usage allowances, and the model itself performs better. Another comment pointed out that the current situation between OpenAI and Anthropic has swapped from earlier this year in April-May, with the community experiencing a "mass migration from Codex to Claude."

Users also had many complaints about the announcement itself: good news is brief and easy to understand, while bad news is written in thousands of words yet logically confusing. One person directly called the original post author a "master of manipulating user emotions," calling it an "extremely insulting post," and said reading it further reduced their expectations for DevDay.

Original thread
02

Zhihu

3 selected conversations

Xiaomi's LLM Head Luofuli Promoted to Level 22: Level Significance, Xiaomi AI Strategy and Public Debate

LatePost learned on September 22 that in Xiaomi's promotion list, Luofuli, the 31-year-old head of the LLM team, was promoted to Level 22. People close to Xiaomi pointed out that Level 22 is already the highest level in Xiaomi's rank system, with further promotions mainly reflected in job title changes rather than numerical growth.

The high-scoring answer analyzes this promotion in the context of Xiaomi's AI architecture adjustment: this July, the technical capabilities of Xiaoai were split into three parts—the foundational model went to Luofuli's MiMo team, edge went to the Mobile and Automotive OS team, and cloud engineering went to Luan Jian. This answer argues that the foundational model is the "brain," edge is the "hands and feet," and cloud is the "pipeline," with Luofuli managing this core layer of the foundational model. In conjunction with MiMo's recent release of both Pro and Flash dual models—the former focusing on complex long tasks, the latter on high-frequency calls and large-scale workflows—this answer argues that she "managed to produce something at this juncture that can still be applied to actual business," making her promotion unsurprising. The answer also points out that Xiaomi is using AI as a bottom-layer strategic layout, with phones, cars, and ecosystem all relying on LLM capabilities. "Whoever controls the foundational model controls the entry point," and Level 22 serves both as an internal incentive signal and an external declaration of Xiaomi's intent to compete for AI talent.

Another answer takes a different angle, starting from Xiaomi's overall technical path, arguing that its LLM follows a route of "competing on efficiency rather than intensity," constrained by hardware and funding from competing in compute with top-tier vendors, therefore persisting with lightweighting and efficiency optimization, "and indeed found a path." Comments added that according to public information, this team consists of only about a hundred people, as exploratory R&D teams are typically maintained at 10 to 30 people, not the same concept as large-scale engineering teams.

Luofuli's promotion to Level 22 this time left many people's first reaction as how young she is at 31.

This answer also records dissenting voices that appeared in the Tieba community, with subsequent comments offering numerous rebuttals. Some users sarcastically referred to "Huawei's 'Qianwen' AI" and pointed out that "America and Xiaomi are too terrible, forcing Huawei to plagiarize Alibaba's Qianwen," using absurd irony to question the original poster's logic. Other comments summarized by saying "Before Huawei develops something, using it makes you a spy; after Huawei develops it, not using it makes you a spy," describing the logical predicament present in the related discussions.

Additionally, some comments mentioned a certain model's performance gap between evaluation platforms and actual experience, saying "ranked first in benchmarks, yet can't do anything right," tying it to Xiaomi phone reviews, while also stating that they themselves use Xiaomi products.

Original thread

Claude Sonnet 5.5 Benchmark Scores Soar, Pricing Sparks Heated Debate

Released on September 28, Claude Sonnet 5.5 showed significant jumps across multiple benchmarks. Terminal-Bench surged from 10.3% to 70.6%, surpassing Opus 5.5's 66.4%. Terminal-Bench-Science also entered the top five. Coding ability comparison shows the score rising from 0.7% to 46.1%. Some users noticed that the new version introduces a safeguard mechanism for high-risk tasks—when encountering "high-risk" requests, it falls back to Sonnet 5. Based on this, they speculate that all Anthropic products are now equipped with fallback. This release happened exactly one day before OpenAI DevDay, and was playfully called a "rubbing it in" style of timing.

On pricing, one user did the detailed math: Sonnet 5.5 API unit price is only one-fifth of Astra's, but the per-task consumption is more than double that of Astra. By efficiency, the gap exceeds tenfold. By subscription benefits, Sonnet quota is close to double that of Opus, and the $20 package equates to approximately $3,000 worth of API usage, offering notable value-for-money. This user believes Sonnet 5.5 is suitable for maintaining existing projects and daily small iterations, but not yet sufficient for building large commercial projects from scratch. However, the $20 tier subscription is enough to handle quite a few scenarios. Domestic users report that fewer and fewer people can access Claude, so anyone routing requests to Claude might feel an improvement in experience.

Skepticism also appeared in the discussion: some questioned whether anyone actually uses Sonnet for daily tasks, suspecting that promotional soft articles exist in the Chinese community. Others pointed out that running an AI relay service is the most profitable, suggesting positive reviews could be spontaneous promotion from believers. There were also opinions that Sonnet still falls short of Astra Pro in mathematical research. In anecdotal statistics, over 90% of paying users doing mathematical research choose GPT.

Regarding vendor strategy, some users compared Anthropic to TSMC—the lab may have newer versions but commercializes the current one first, with long release cycles. OpenAI, on the other hand, is like playing all its cards each time, products often feeling rushed out with bugs unfixed. In response, some users argued they have never heard of Claude being dumbed down, whereas OpenAI frequently swaps models quietly. Regarding whether OpenAI still has advantages, comments were divided: some believe OpenAI has fallen behind in all aspects, while others pointed out that OpenAI has its own moat, and both companies have their own strengths. It's just that OpenAI can no longer swagger like it used to.

Original thread

Behind Anthropic's Sprint to a $2 Trillion Valuation: The Real Game Between Revenue Growth and Operating Losses

Anthropic's IPO filing shows 2025 revenue of $4.6 billion, with plans to target a $2 trillion valuation in 2026, implying a $65 billion revenue run rate. However, operating losses reached $8.06 billion, with compute and infrastructure spending exceeding $7.3 billion—1.6 times revenue—and future committed investments as high as $518 billion. Revenue is highly concentrated, with roughly a quarter coming from two clients; Meta contributed over $5 billion. On governance, the seven co-founders control 50.1% of voting rights through Class F shares. The filing dedicates 80 pages of risk factors against 48 pages of business description, a style described as "Anthropic's style," and explicitly states it will not develop image or video generation models.

Whether the valuation is reasonable has sparked disagreement. Some commentators argue that on an ARR basis Anthropic is relatively cheap—DeepSeek's ARR is only around $800 million, a gap of tens of times—but rebuttals point out that DeepSeek only has that revenue because of external constraints, so the two are not simply comparable. Others subtly mocked some companies for "actually thinking their business is worth that much in costs." Others observed that "the more one warns about AI products posing extinction-level risks, the more capitalists believe in their capabilities."

Anthropic's accusation that Chinese models "distilled" Claude has sparked heated debate. Some view this as being "driven to having no tricks left"—Chinese open-source models have already surpassed some U.S. closed-source models in popularity, "distillation isn't a big deal, U.S. companies distill each other too," implying the security narrative is meant to slow competitors' R&D progress. But this claim has been questioned: the so-called 50% figure comes only from OpenRouter platform token share statistics (~60%), not user numbers, with actual call ratios between 30% and 46%.

On the enterprise application side, users have shared their experiences: in early 2026 using only Claude, adding OpenAI after GPT 5.5's release, and by August switching to self-deployed domestic open-source models, because "$500–$1,000 a day is easy to burn through," now afraid to use it aggressively, with even JPMorgan reportedly "only giving each person $2,000 a month." Large companies generally do not allow employees to use external AI—there's the cost, the contribution of training data, and the leakage risk.

Original thread
03

Hacker News

3 selected conversations

OpenAI Releases GPT 6.1 Sol: API Prices Slashed, Cached Input as Low as $0.10/M

OpenAI released GPT 6.1 Sol at Developer Day, positioned close to the top-tier Astra, with pricing at about one-fifth of the latter. The core highlight is the cached input price: just $0.10 per million tokens, a 95% reduction from standard input, and another 50% reduction from GPT-6 Sol's cached price. Some users calculated that combined with the low cache rate, the available volume on Codex will increase significantly.

Multiple users have expressed dissatisfaction with the GPT-6 family. Sol 6 is being criticized as a "severe regression" compared to Sol 5.6, with Luna facing similar issues; even the top-tier Astra proves unreliable for programming tasks, despite strong vision capabilities but frequent errors. One user admitted to promoting OpenAI/Codex models for the past half-year but is now "very disappointed" with OpenAI and skeptical about whether 6.1 can improve.

A comment uncovered that a model called Astra-Minor appeared in pre-release documents days before launch, speculating that Sol 6.1 was urgently renamed from that model. Given Sol 6's mediocre performance and Opus 5.5 exceeding expectations, the team likely made last-minute adjustments—otherwise it's hard to explain why a new version emerged just days after Sol 6's release.

Some users prioritize cost-performance over raw capability. One person noted spending less than $200 on Deepseek this year, finding the monthly $200 or $500 subscription hard to justify without clear usage scenarios; Deepseek is faster, the intelligence gap is minimal, and the price is cheap enough that users stop worrying about consumption volume. Another user stated they are willing to trail the "frontier" by roughly six months in exchange for better value, and doesn't care which "mysterious government" their data belongs to.

On the same day, OpenAI launched a $500/month Pro high-end subscription providing maximum usage limits and "Ultrafast" features. However, the $200 Pro subscription benefits were simultaneously reduced: usage in Codex and Work dropped from 20x to 10x of Plus tier, and GPT-6 Pro's weekly message limit in ChatGPT was cut from 200 to 100. Some in the comments directly exclaimed "Well, well." Others view the price war as a signal of AI model commodification, arguing that big labs are sliding into "race to the bottom."

At the specific use case level, GPT 6-Astra's 3D modeling capability in Blender is still considered irreplaceable, while Opus 5.5 is stronger in programming. Some comments question the phenomenon where higher effort levels in benchmarks yield lower scores—for example, GPT-6.1 Sol High scored 75.2% on DeepSWE while the pricier XHigh only scored 71.9%—and question whether the tests were only run once.

Original thread

OpenAI's Resident Agent Dots Sparks Debate: Lock-in Risks and Unclear Positioning

The news of OpenAI launching its resident agent Dots sparked discussion on Hacker News, with multiple comments placing Dots alongside Codex, ChatGPT Work, Meta's Muse, Grok Bots and Instinct, arguing that the boundaries between these similar products are becoming increasingly blurred—some comments noted that the ideal form should be a remote agent with long-term memory in a cloud sandbox, with no need for each to go its own way.

However, concerns about lock-in ran throughout the discussion. Some pointed out that switching between models is relatively easy, whereas resident agents, due to integration with platforms like social accounts, email and work records, carry extremely high migration costs—'they are essentially your cloud computer.' One comment further speculated that closed-source model companies are eager to build abstraction layers to restrict users from directly accessing models. Another comment targeted AI companies' software-building capabilities, calling them 'good at models, weak at software,' citing Anthropic as an example to illustrate how a user base built on subscription models can easily churn after product strategy missteps.

Some users held reservations about the product design itself. One comment linked the cartoonish style to the severity of product issues, arguing that 'cute design often signals invasiveness.' Others criticized a common flaw among leading AI companies: the gap between what is demonstrated at launch and what the actual product looks like half a year later, with budget cuts leading to severe feature reduction. On the privacy front, one user explicitly stated they would not try it unless they could guarantee it 'only serves their own interests,' including 'never sharing my secrets.'

I've never had that happen with Grok Bot bots because I have one for my biz admin and one for my personal admin.

Positive feedback came from users who have been long-term Grok Bots users. This user stated that multiple resident agents with separated responsibilities can effectively prevent context from being polluted by irrelevant memories, and can handle multi-task workflows without human supervision. But they admitted that after use, inference costs increased by 2 to 3 times, and completion time extended from approximately 10 minutes with local prompting to approximately 1 hour. Other users expressed confusion: their work was constrained by manual approval stages, the agent couldn't complete tasks independently, and at night they couldn't find any tasks to assign. Another comment mentioned that Dots is explicitly not available to users in the European Economic Area, Switzerland, and the UK.

Original thread

Sonnet 5.5 Reviews Focus on Pricing Controversy and Agent Tool Gaps

The release of Sonnet 5.5 sparked heated debate, with pricing as the central focus. Multiple commenters considered the price too high, citing comparison data from aibenchy.com that its performance is roughly on par with luna yet costs 14 times more; other commenters pointed out that, according to Anthropic's official benchmarks, setting the thinking level to medium or above causes costs to quickly approach or even exceed Opus 5.5. Supporters noted that the model has become the new default option for the free tier, giving users of Claude Web's free version access to frontier model capabilities.

In specific scenarios, the value of upgrading received validation. One developer benchmarked using Bakeoff (which uses real bugs and features, scored by held-out tests), finding that Sonnet 5.5 is 5.5x faster than Sonnet 5, costs reduced to a quarter, call rounds cut to a third, and supports batch reading and completing edits and tests in a single call.

Regarding thinking level settings, a commenter advised free users to set it to Medium rather than Max, characterizing Medium as "more patient and capable of completing tasks," while Max is "prone to irritability and verbosity," with Medium mode showing noticeably slower token consumption (though this experience is limited to September 2026). In professional use, another commenter raised a practical concern for text classification scenarios: smaller models are faster but miss edge cases, the accuracy vs. latency trade-off always exists, and questioned whether Sonnet 5.5 truly improves the false positive rate.

Multi-model collaborative workflows also sparked discussion. One person described assigning 80% of tasks to Sonnet 5.5, with Opus 5.5 handling the remaining 20% for refinement, and Gemini 3.1 Pro serving as researcher and design verifier. Another commenter noted that Entropic has achieved breakthroughs in certain areas, and raised an open question: whether tools exist that allow Anthropic models to call sub-agents or dynamic workflows from other providers, thereby creating cross-provider agent swarms.

Some commenters observed OpenAI's recent aggressive moves, including shelving Astra 6.1 and filing an S-1 for an IPO. Other commenters speculated about Anthropic's strategic intentions, suggesting that Sonnet 5.5's release may be the company's attempt to challenge OpenAI's position in the public market under pressure from being excluded by the US government. However, all of the above are personal conjectures and not community consensus.

Original thread
04

Xiaohongshu

3 selected conversations

Fei-Fei Li Joins AMD: Can Spatial Intelligence Surpass the Large Language Model Debate

The original post reported that AMD acquired World Labs in an all-stock transaction of approximately 8.2 billion USD, with Fei-Fei Li set to join AMD as Executive Vice President and Chief Scientist upon completion, reporting to CEO Lisa Su. The core argument of the original post was not the amount of the transaction but rather Li's judgment on AI development direction: "The world is not made of text." She argued that while large language models like ChatGPT, Claude, and Gemini have made remarkable progress in text processing and reasoning, objects in the real world have distance, depth, weight, and spatial relationships. A robot reaching for a cup not only needs to "know what a cup is," but also understand its position, movement, and collision. In 2024, she founded World Labs, betting on "Spatial Intelligence"—defining it as a foundational model direction that transcends LLMs. The original post noted that World Labs had already collaborated deeply with AMD last year, conducting model training and inference optimization on AMD GPUs, and that when AI truly moves toward robotics, autonomous driving, design simulation, and the entire physical world, chips, software, foundational models, and applications need to be "fully integrated." Based on this, the original post suggested that AI competition may be entering a new phase: the first phase enabled machines to understand text, the second phase enabled machines to perceive images and video, and the next phase will enable machines to truly understand the physical world.

The comments section showed clear disagreement. Some users directly expressed skepticism about scientists and AI development, saying "I've lost my reverence for scientists now," believing that people who changed the world in recent years "look back and have basically made the world worse," and stating "I still can't see AI benefiting humanity in any aspect," and asserted that "technology doesn't necessarily drive progress, but it definitely drives involution." Other users believed that the real controversy behind this acquisition is whether continuing to scale up large language models can lead to AGI, or whether AI must truly understand the physical world, saying "this may be the most important route debate in AI in the coming years." Still other users pointed out from an application perspective that if so much hot money is flooding into AI applications in the biomedical industry, the new drug development cycle will be significantly shortened, costs will drop, and understanding of pathophysiology will also progress faster, "the problem is that these things can't attract the attention of investors and the public."

Original thread

Users in discussion forums share results and limitations of converting Yanyun Shisheng Voice screenshots into CCD retro camera style using GPT

A user in a discussion forum shared an attempt to use GPT to process screenshots from the game Yanyun Shisheng Voice into a retro CCD camera effect. The original post included an English prompt asking to generate photos with nostalgic CCD texture, direct flash, slight overexposure, slightly cool magenta tones, visible noise, soft blur, and the look of an early digital camera. In the comments, some users mentioned that Teacher G's character customization effect was very realistic, while others said the processed results looked like genuine photos casually taken by someone walking down the street.

However, this prompt was presented as an image and couldn't be directly copied, which drew complaints from other users. Another user dug up a close-up screenshot of General Wang Qing from the game to try it, describing the generated result as "totally amazing". The remaining comments were mostly emoji or brief exclamations, with the overall atmosphere being predominantly positive feedback.

Turn this screenshot into a realistic photo with a nostalgic CCD aesthetic : candid snapshot , direct flash, slight overexposure, cool-magenta tint , visible noise, soft blur, and early-2000 s digital camera vibes.

Original thread

Anthropic Accused of Implementing "Order-Grabbing" Workflow, Comment Section Questions Source Reliability and Draws Parallels to Domestic Tech Giants

A post titled "Anthropic's Internal Toxic-to-the-Core Workflow" sparked discussion. The original post described Anthropic's internal working methods as similar to the order-grabbing mechanism on food delivery platforms, and commented that "compared to this order-grabbing workflow of 'A-beast,' 996 is truly a blessing," while also expressing concerns about domestic tech giants potentially adopting this model.

Among the comments, someone directly likened the order-grabbing mechanism to "Meituan Waimai's order-grabbing hall," someone called it "real-person async RL," and others mentioned "Google started this years ago," implying that such speed-as-core-metric work models are not new. One comment asked: "This is just the speed dimension of evaluation criteria—what about the quality dimension?" questioning whether a single speed metric is reasonable.

Many comments continued with sarcastic banter: "Then I'll just poison things—ask me and I'll call it adversarial testing," "Holding tight, waiting to grab and do someone else's work once they're almost done," "@grok check for me every 30 seconds for sniping opportunities," "/goal sabotage everyone's work." These replies all jokingly suggested that if the workflow truly centers on order-grabbing, the game-playing behavior among employees could head in a destructive direction.

At the same time, some comments questioned the credibility of the post. One comment pointed out "This doesn't look very credible—if credit attribution is this unclear, it would have fallen apart long ago," arguing that such a management model would be hard to sustain with unclear contribution attribution; another comment labeled it "source: trust me bro," joking about the original post's lack of reliable sources. Someone also asked "Did the founder get this inspiration from his early days at Baidu?" but received no response.

Original thread
About this issue

From AI discussions collected that day with new comments, up to three top discussions are selected from each community. Duplicates on the same topic are removed, and results are sorted by the number of comments collected that day; this is not a platform-wide ranking.

Discussion window: 2026-09-29 (UTC); source material extracted on 2026-09-30 (UTC), which may include subsequent edits. User self-reports, predictions, and paraphrased content retain attribution; non-English excerpts in this English edition are translated to English, while English excerpts remain in the original language.

Explore previous issues →