Doubao 5.1 Multimodal Fails Historical Figure Test as Qwen, DeepSeek Pass

A Zhihu comparison test reveals Doubao 5.1 misidentifying notable figures as unrelated roles while competing Chinese AI models analyze historical imagery accurately.

Hacker News 2 · Reddit 7 · Zhihu 5 480 covered discussions 7 source-linked evidence passages

The day in brief

A viral Zhihu comparison test has exposed stark performance differences among Chinese AI models in analyzing historical imagery, with Doubao 5.1 notably misidentifying prominent figures—misreading terms like 常凯申 and 少帅 as occupations—while Qwen and DeepSeek handled the same images with varying degrees of accuracy. Meanwhile, new midsize Qwen 3.8 and Ornith 1.5 models have arrived with improved quantization options, and the community continues exploring unconventional AI workflows for book shopping and travel planning. Claude Code's tendency to overestimate task duration by days while delivering in minutes has also sparked discussion, alongside an ongoing debate over whether free-tier DeepSeek can substitute for ChatGPT Pro's $200/month subscription.

Model experience tracking

Doubao 5.1 Multimodal Failure Exposes Gaps in Historical Figure Recognition

The Incident

  • Doubao 5.1 misidentified a photograph of historical figures 常凯申 and 少帅 as depicting 轿夫 (sedan chair carriers), incorrectly describing them as such. The model also failed at basic counting, identifying four people in the image as five.
  • In contrast, Qwen correctly identified the figures in the same photograph without error.
  • DeepSeek V4 Flash refused to process the image entirely, declining to provide any analysis.

Online Response

  • A Zhihu post comparing model responses to the identical image gained significant traction, accumulating 44 total engagements with a peak score of 36.
  • The post author drew a clear conclusion for readers, recommending that users avoid Doubao for multimodal tasks, writing: 'Now you know which one to use.'
  • A commenter added levity to the discussion, joking that DeepSeek 'pretends not to be multimodal' but 'runs away first' upon encountering sensitive images.

Key Takeaways

  • Community consensus from the discussion points to Qwen as the preferred model for image analysis tasks requiring accurate figure identification.
  • The test highlights a specific failure mode for Doubao 5.1 on historical figure recognition tasks, demonstrating limitations in its multimodal capabilities.

Applicable Scenarios

  • Comparative evaluation of multimodal models on historical figure identification tasks.
  • Testing model response to images containing Chinese historical figures.

Practical Value

  • The incident provides concrete evidence of performance variance across models when processing identical input images.
  • It illustrates how different models handle sensitive or semi-sensitive historical imagery, ranging from confident misidentification to outright refusal.

The Test Prompt

  • A Zhihu user submitted an image of 常凯申 and 少帅 for analysis, requesting identification of the figures in the photograph.

Prompt Analysis

  • The prompt tested whether models could correctly identify historical figures in a photograph and accurately describe their roles and appearances.
  • The test evaluated both identification accuracy and counting precision across multiple models simultaneously.

Community evidence

Looking at Yuanbao: no problem. Looking at Afu: completely unrecognized. Now you know which one to use.

DeepSeek vs. ChatGPT Pro: Why 99% Choose Free—And the 1% Who Cannot

The Core Debate

  • Chinese AI communities are debating whether free-tier DeepSeek can match the capabilities of ChatGPT Pro, which costs $200 per month. The discussion examines the distinct advantages of each platform.
  • The GPT web version offers nearly unlimited chat quota on the Pro tier, compared to a 3,000 messages per week limit on Plus and Team plans. GPT Pro also provides access to a cloud virtual machine through GPT Work, equipped with a 9-core EPYC processor and 15 to 22 GB of RAM.
  • Users report that GPT Image 2 delivers what some describe as the strongest image generation in the industry, with a nearly unlimited quota. GPT Pro also includes scheduled monitoring and alert notification features that allow users to track specific topics over time.
  • DeepSeek is widely praised for its Chinese text style mimicry, with users asserting that no current model can match its precision in capturing nuanced Chinese language patterns. In a practical comparison, DeepSeek completed a complex VBA code-writing task in just 5 conversation turns, while GPT produced only the first page after 10 turns.
  • One developer reports using DeepSeek for production line operations handling 100 concurrent requests, while deploying GPT for content moderation due to stability differences under heavy loads.

Community Perspectives

  • A high-engagement answer with a score of 97 states that 99% of users do not need ChatGPT Pro at $200 per month—they simply do not use enough of its capabilities. However, the remaining 1% who require extreme research, high concurrency, fast response times, code generation and analysis, and advanced multimodal features find the gap between the two services 'completely on a different level.'
  • Community members hold sharply differing views on GPT image quality. One user firmly states that GPT Image 2 is the industry's strongest image generator, with no viable competition, while another explicitly rejects this assessment.
  • Users note that the GPT web version rarely encounters hallucinations, a frustration that some have experienced with Gemini. One commenter argues that GPT web version produces philosophical responses of exceptional quality, exceeding what even specialized professionals typically write.
  • The broader community sentiment suggests a complementary approach: using DeepSeek for production and creative tasks while deploying GPT for moderation and monitoring roles.

Key Takeaways

  • For 99% of typical users, free DeepSeek provides sufficient capabilities for their needs, making the $200 monthly ChatGPT Pro subscription difficult to justify on a cost-benefit basis.
  • For users with extreme research requirements, high concurrency demands, and advanced multimodal tasks, power users report that GPT Pro remains unmatched and irreplaceable.
  • DeepSeek offers distinct advantages in Chinese language tasks, concurrent load stability, and style mimicry that make it the preferred choice for specific use cases.
  • GPT Pro's value proposition centers on its unlimited quotas, cloud VM access, and top-tier image generation capabilities rather than raw conversational quality alone.

Practical Applications

  • Complex academic research requiring high message volume and extended reasoning sessions over long conversations benefits significantly from GPT Pro's generous quota.
  • Production systems requiring stable high-concurrency API access find DeepSeek more reliable under load, while GPT handles complementary moderation tasks.
  • Professional Chinese language content creation requiring precise style mimicry benefits from DeepSeek's unmatched capabilities in this area.
  • Automated monitoring and scheduled task execution leverage GPT Pro's alert notification features to track specific topics over time.
  • Large-scale image generation workflows requiring both quality and volume find GPT Image 2's capabilities unmatched by current alternatives.

Value Assessment

  • DeepSeek provides cost-effective access for general users with Chinese language and style-specific needs, delivering strong performance without a subscription fee.
  • GPT Pro justifies its premium price for users requiring unlimited quotas, cloud compute resources, and advanced multimodal capabilities that free or lower-tier services cannot provide.
  • Combined usage patterns are emerging in the community, where users deploy each service for tasks aligned with its particular strengths rather than relying exclusively on a single platform.

Discussion Framing

  • For typical users: The central question of whether free DeepSeek can match $200/month ChatGPT reveals that 99% do not need the Pro tier, though 1% of power users find the capability gap irreplaceable for extreme research, concurrency, and multimodal work.
  • For GPT web advantages: GPT Pro offers nearly unlimited chat quota, Gmail and GitHub integrations, a GPT Work cloud VM with 9-core EPYC processor and 15 to 22 GB of RAM, top-tier GPT Image 2 generation, and scheduled monitoring features.
  • For DeepSeek strengths: Chinese text style mimicry unmatched by any current model, stable performance at 100+ concurrent requests, and the ability to complete complex VBA tasks in 5 turns versus GPT's partial results in 10 turns.

Discussion Dynamics

  • The debate centers on whether free-tier capabilities can substitute for $200 monthly subscriptions, reflecting strong cost-consciousness within Chinese AI communities as users evaluate their actual needs against premium pricing.
  • The discussion reveals clear model specialization patterns rather than universal superiority of either platform, with each service excelling in different domains.
  • Community sentiment reflects practical triage: deploying DeepSeek for routine and production tasks while reserving GPT Pro for mission-critical high-volume work.
  • Image generation quality remains a contested topic with polarized community opinions, particularly regarding GPT Image 2's claimed industry leadership position.

Community evidence

For 99% of people, the $200 ChatGPT is completely unnecessary, but for the 1% who need it, ChatGPT at $200 is beyond what DeepSeek can reach, especially in extreme scientific research applications, high concurrency states, response speed, code generation and analysis, and multimodal capabilities.

GLM-5.3 Matches Kimi K3 on Artificial Analysis Benchmarks Through Efficient Post-Training

Performance Achievement

  • GLM-5.3 has achieved parity with Kimi K3 on the Artificial Analysis benchmark index, representing a significant milestone for the Chinese AI developer. The achievement is notable because it was accomplished using the same base model, same architecture, same total parameters, and same activation parameters as the previous version, with approximately one month of post-training and reinforcement learning—without any changes to the foundational model itself.
  • This result positions GLM-5.3 within the competitive mid-high capability tier on the Artificial Analysis index, placing it alongside established models such as DeepSeek V4 Flash and GPT-5.6 Luna. The shorter thinking time demonstrated by GLM-5.3 compared to these competitors suggests a different approach to training efficiency, potentially indicating that parameter count alone may not be the determining factor for coding performance.
  • Additionally, GLM has integrated Anthropic-style network security capabilities designed for enterprise deployment, suggesting a strategic focus on the government and enterprise market segment where such features are increasingly valued.

Practical Applications

  • GLM-5.3 is particularly well-suited for coding tasks that benefit from efficient reasoning without requiring extended chain-of-thought processes. The achievement of competitive results at the 700B parameter scale suggests that certain coding workloads may not require the excessive computational overhead previously assumed necessary for high performance.
  • Enterprise deployments requiring integrated network security capabilities represent another key use case. GLM's partnership approach with security vendors positions the model as a solution for organizations seeking both advanced AI capabilities and built-in security features for sensitive environments.

Comparison Framework

  • The prompt requests a direct benchmark comparison of thinking time between GLM-5.3 and DeepSeek V4 Flash on coding tasks. This comparison is reproducible as it asks for measurable performance metrics that can be tested on the Artificial Analysis platform or via OpenCode with consistent coding tasks. The focus on thinking time rather than raw output quality reflects the community's growing interest in understanding the efficiency characteristics of different model architectures and training approaches.

Benchmark Comparison Prompt

  • Compare thinking time between GLM-5.3 and DeepSeek V4 Flash on coding tasks.

Key Insights

  • The 700B parameter scale may be sufficient for coding tasks without requiring excessive chain-of-thought reasoning, with remaining performance gaps potentially depending more on access to high-quality executable environment data for training rather than further parameter scaling.
  • Post-training efficiency can achieve competitive results without base model changes, suggesting that targeted reinforcement learning and optimization efforts may offer a cost-effective path to improved performance for models already operating at sufficient scale.

User Responses

  • Some developers are switching from Claude to GLM-5.3 via OpenCode, noting that despite the high switching costs associated with Claude Code plugins, the transition has been worthwhile. Users specifically value the shorter thinking time, as GLM-5.3 achieves results without requiring prolonged chain-of-thought processing.
  • However, concerns have emerged regarding the trajectory of proprietary models becoming increasingly opaque. One developer noted that models like Codex now encrypt agent-to-agent communications, creating 'black box' scenarios where users have limited visibility into what subagents are instructed to do or what they report back, raising questions about transparency and trust in commercial AI systems.

Value Proposition

  • GLM-5.3 achieves competitive mid-high tier performance on the Artificial Analysis index through an efficient training approach, demonstrating that meaningful performance improvements can be realized through targeted post-training optimization without the resource requirements of training a new base model from scratch.
  • The shorter thinking time achieved by GLM-5.3 suggests that efficient training rather than parameter reliance may be a viable strategy for models at the 700B scale, potentially offering better latency characteristics for interactive coding applications where response time matters.

Community evidence

Models are not necessarily stronger with more parameters; you also need to consider training data, training duration, inference costs, and whether it's a dense model or MoE.

Claude Opus 5.0's Verbosity Paradox: Superior Code Generation Accompanied by Overblown Writing

The Verbosity Phenomenon

  • Claude Opus 5.0 produces output that expresses simple concepts with great complexity and bombast, according to user reports.
  • Pre-emptive hedging appears in Claude Opus 5.0 output as a substitute for working memory; the model writes comments such as 'this no longer does an O(n^2) read over all rows' because the order of operations involves generating the quadratic version first before the user's prompt steers it toward linear, yet the explanatory comment persists as a record to its "amnesiac future self."
  • A research paper (arxiv:2507.02618) on strategic behavior in iterated prisoner's dilemma found distinctive persistent "strategic fingerprints" across LLM families: Google's Gemini models were strategically ruthless, exploiting cooperative opponents and retaliating against defectors, while OpenAI's models remained highly cooperative, a trait that proved catastrophic in hostile environments, and Anthropic's Claude was more cooperative still but outperformed OpenAI head-to-head.
  • Despite anti-sycophancy tuning, Claude Opus 5.0 reportedly maintains its fundamental cooperative strategic pattern and continues finding things to nitpick regardless of style guide modifications.

User Observations

  • Users report that while Claude Opus 5.0's output programs are superior to prior iterations for their use-case, the writing style exhibits an "amusing degree of bombast."
  • Friends share examples of Claude producing verbose outputs like "to be honest, it sounds like you" as a humorous contrast to its typically overcomplicated style.
  • Observers note the personality of the underlying model persists even when surface tone is altered, referencing difficulty with Gemini models in RPG character roles and relating this to Claude's persistent strategic behavior.

Practical Applications

  • Programming and code generation tasks reportedly benefit from Claude Opus 5.0 improvements over prior iterations.
  • Verbose writing style becomes a friction point for use cases prioritizing conciseness.

Key Insights

  • Style guide modifications do not fully override Claude Opus 5.0's tendency to find something to nitpick; the fundamental strategic behavior persists beneath surface-level tone changes.
  • Pre-emptive hedging in generated code comments reflects working memory limitations during the token generation process rather than intentional documentation.

Strategic Value

  • Understanding persistent "strategic fingerprints" across LLM families helps set appropriate expectations for Claude's cooperative behavior and tendency toward preemptive hedging.
  • Awareness that hedging comments serve as working memory substitutes can inform prompt engineering strategies to reduce unnecessary verbosity.

Core Question

  • What tuning resulted in simple concepts being expressed with great complexity?

Question Context

  • This prompt reproduces the user's direct question from the source material, seeking the root cause of Claude Opus 5.0's verbose output style.

Community evidence

The personality of the underlying model persists even if you're able to alter surface tone enough for your needs.

Ecosystem and open models

New midsize Qwen 3.8 and Ornith 1.5 models reshape local AI landscape with improved quantization options

Community response to new model releases

  • Users express excitement about the Ornith 1.5 family, with one joking 'We have Q3.8 35B at home' when comparing the new releases to expectations for an official Qwen 3.8 35B model, a post that garnered 37 mentions and 80 engagement.
  • A community member reports that Qwen 3.8 27B struggles with agentic coding tasks, noting the model runs excessive tokens, loops, or fails to complete tasks; this user runs the Q6_K quant on 2x3090Ti with LM Studio, with the report accumulating 40 mentions.
  • Users observe that Qwen 3.8 27B lags behind DeepSeek V4 Flash without thinking mode but performs impressively with thinking enabled; the community wants a lower active parameter model that supports thinking mode at good speed.
  • Community members express enthusiasm about combining 35B A3B models with thinking-enabled larger models; users suggest using 35B models for speed while enabling thinking for quality, creating a flexible workflow.

Key practical insights for users

  • Unsloth Dynamic v3 1-bit quantization achieves 77% accuracy retention, enabling Qwen3.8-27B to run on systems with only 8GB RAM, dramatically lowering hardware barriers for local deployment.
  • MTP (Multi-Token Prediction) remains available as a separate component in Unsloth quants for users who want to opt in, allowing flexibility without forcing the feature on all users.
  • Community notes suggest running 35B models for speed and enabling thinking mode for quality, providing a practical framework for model selection based on task requirements.

Value provided by new releases

  • The Ornith 1.5 family provides new benchmark-competitive alternatives in 9B, 35B, and 397B sizes with full GGUF support, offering options across different hardware capabilities and use cases.
  • Unsloth Dynamic v3 offers an improved accuracy-size tradeoff for Qwen3.8-27B quantization, with new versions delivering 10% higher accuracy for the same model size and outperforming alternatives by more than 10% on Div-300 and KLD benchmarks.

Original user request

  • After extensive testing, Qwen 3.8 27B lags nicely behind DeepSeek V4 Flash without thinking. With thinking on, it is seriously impressive. The point being, we really need a high total, lower active parameter model that we can engage thinking on that will be fast enough.

Analysis of user needs

  • The user reports that Qwen 3.8 27B underperforms DeepSeek V4 Flash without thinking mode but becomes competitive when thinking mode is enabled, indicating a clear need for smaller active parameter models that support efficient thinking mode to balance speed and quality.

Recommended applications

  • 35B parameter models are recommended for faster inference speeds, making them suitable for real-time applications and workflows where latency matters more than maximum quality.
  • Thinking-enabled mode is recommended when quality is prioritized over speed, allowing models to spend additional compute on reasoning tasks where accuracy is critical.

Key developments this week

  • Ornith AI released the Ornith 1.5 family with three model sizes: 9B, 35B-A3B, and 397B, all available with GGUF variants on Hugging Face.
  • Ornith claims that Ornith 1.5-35B surpasses Q3.6 27B on Terminal Bench 2.1, SWE Bench Pro, and DeepSWE benchmarks, and overtakes Q3.8 27B on the NL2Repo benchmark.
  • Ornith claims that Ornith 1.5-397B surpasses DeepSeek V4 Flash 0731 and Opus 4.8 according to their benchmark results.
  • Unsloth released Qwen3.8-27B Dynamic v3 GGUFs claiming 10% higher accuracy for the same size, outperforming others by more than 10% on Div-300 and KLD benchmarks.
  • Unsloth Dynamic v3 released 1-bit quants for Qwen3.8-27B retaining 77% accuracy, runnable on systems with just 8GB RAM, significantly expanding accessibility for lower-end hardware.
  • Unsloth uses post-training quantization without training on imatrix calibration dataset and does not use QAT or QAD, relying on standard post-training techniques.
  • MTP was removed from Unsloth quants after community feedback and uploaded as a separate component, giving users the choice to add it if desired.
  • A Qwen community manager announced a new midsize Qwen 3.8 model in the 35B-100B range coming next week, with no early access program due to scheduling constraints.

Community evidence

I compared the values with Q3.8 27B: What they claim is that their 35B surpasses Q3.6 27B (Terminal Bench 2.1, SWE Bench Pro, DeepSWE) and even overtake Q3.8 27B in NL2Repo benchmark.

Real use and unexpected gains

Reddit Community Shares Unconventional AI Workflows That Save Hours on Book Shopping and Travel Planning

The Thread That Sparked Unconventional Workflows

  • Reddit users shared unconventional personal AI workflows under a post asking about "weirdly specific tasks" that save hours of time.
  • One user described uploading a file of their favorite and least favorite books—including genre preferences, topics, and prose style dislikes—to create a personalized book-shopping assistant for thrift store visits.
  • The same user reported asking the AI to evaluate books encountered in stores; the AI researches online reviews, retrieves Goodreads ratings and recurring criticisms, and determines fit based on the user's preference file.
  • Another user described using AI to plan multi-attraction travel itineraries from a single base city, with geographic clustering of points of interest.
  • The travel workflow included specifying activity types per day (e.g., one chill and one active) and requesting restaurant recommendations near scheduled visit points.

The Original Reddit Prompt

  • What is a weirdly specific task you use ChatGPT for that actually saves you hours? (Not typical stuff like writing basic emails, coding boilerplates, or summarising long PDFs).

Why This Prompt Surfaced Unconventional Use Cases

  • The prompt specifically excludes common productivity tasks to surface unconventional, niche applications.
  • Scoping to "weirdly specific" tasks invites shareable workflows that others may find applicable or inspiring.
  • The constraint on excluding typical use cases filters for higher-signal examples of real-world AI integration into daily routines.

Two Emerging Use Cases

  • Personalized book shopping assistant for secondhand and thrift store browsing.
  • Travel itinerary planning with geographic clustering and mealtime restaurant recommendations.

Why These Workflows Matter

  • Saves research time by aggregating Goodreads ratings, online reviews, and recurring criticisms in one query rather than manual searching.
  • Reduces friction of map-based planning by automating geographic grouping of attractions and identifying restaurants near visit points.
  • Enables pattern-based recommendations by analyzing what a user likes and dislikes to assess fit for unfamiliar options.

Key Takeaways for AI Users

  • Uploading a curated preference file enables personalized analysis beyond generic recommendations, allowing users to evaluate new options against established taste patterns.
  • Geographic clustering workflows reduce planning friction by automatically grouping nearby attractions, allowing users to focus on specifying activity preferences rather than spatial organization.
  • Combining constraint-setting (e.g., balancing chill vs. active activities) with location-aware recommendations produces more useful itineraries than static list generation.

Community Engagement and Response

  • The original post generated strong engagement with 48 mentions and 42 total interactions across the thread.
  • The book recommendation workflow received higher community endorsement (score 7) compared to the travel itinerary workflow (score 3).
  • The book workflow specifically earned praise for pattern recognition—the AI identifying similarities between the user's favorite and least favorite books to improve recommendations.
  • Community response highlighted time savings as a primary benefit: saving research time on book purchases and eliminating manual clicking and map plotting for travel planning.

Community evidence

I've uploaded a file in my project with a list of my favorite and least favorite books and genre/topics/prose preferences and dislikes, so when I see a book I might like I ask if it's a good fit for me and then it researches reviews online, sends me the Goodreads rating and important reoccurring criticism for the book and then tells me if I should get it or not.

Claude Code's Timeline Paradox: Users Report AI Estimates Days of Work, Then Delivers in 30 Minutes

The Estimation Gap

Claude Code users report the tool routinely overestimates task duration, saying projects will take days or weeks before completing them in 20 to 30 minutes. One widely shared post describes Claude Code estimating a project would require 12 weeks of work and advising the user to hire a full-time developer, prompting the user to respond that hiring Claude Code was precisely the purpose.

Community Theories

  • The phenomenon is treated with humor across developer forums, where users share similar experiences of Claude Code predicting unrealistic timelines before delivering results quickly. Community explanations for the overestimation include: reliance on training data patterns that contain inflated human estimates, deliberate upselling of the tool's perceived complexity, and the model's lack of intuitive understanding of GPU execution speed.

Intended Purpose

  • Claude Code functions as an AI coding assistant that users rely on for task completion rather than scheduling accuracy. The tool is designed to assist with development work, allowing users to delegate coding tasks rather than manage project timelines.

User Guidance

  • Users should calibrate expectations when working with Claude Code: the tool's time estimates may not reflect actual execution speed. Treat duration predictions as rough heuristics rather than reliable scheduling commitments.

Outcome Independence

  • The overestimation itself is not necessarily a functional problem; users complete tasks as intended despite inflated predictions. The end result—working code delivered efficiently—remains unaffected by the discrepancy between predicted and actual duration.

Reproducibility Challenge

  • A prompt demonstrating the behavior cannot be reliably reproduced, as the tool's estimates are non-deterministic. Users observe the pattern retrospectively rather than triggering it through specific inputs.

Observational Nature

  • The phenomenon is observational and reported in retrospect; prompts that trigger specific overestimation are not consistently reproducible. This suggests the discrepancy emerges from underlying model behavior rather than prompt engineering, though the exact mechanism remains unclear.

Community evidence

It's hilarious because it told me this project would take 12 weeks and that I should hire one FTE developer.