en.Wedoany.com Reported - On August 3, Alibaba officially released its next-generation foundation model, Qwen3.8, with a total parameter count of 2.4 trillion, featuring significant improvements in coding and professional office (Cowork) capabilities. In the authoritative third-party ranking Arena released today, Alibaba's Qwen model ranks second only to Anthropic's Claude series, placing its overall performance in the global first tier of large models. Currently, the Qwen3.8 API has been launched on the Qianwen AI platform and integrated into Alibaba's Agent product "Qianwen Office," also released today. Qwen3.8-Max is expected to be open-sourced next week, along with Qwen3.8-27B.
Starting today, developers worldwide can access Qwen3.8 API services through the Qianwen AI platform. Domestically, pricing is 12 RMB per million input tokens and 36 RMB per million output tokens, with implicit cache hits costing only 1.5 RMB. Internationally, input and output prices are just 40% and 24% of Opus 5's, respectively, delivering comparable intelligence at higher cost-effectiveness.

This release features Qwen3.8-Max, the largest and most powerful flagship model in the Qianwen series, which supports visual understanding and a context length of up to 1 million tokens. In terms of model architecture, Qwen3.8 leverages a joint optimization of sparse MoE architecture and hybrid attention mechanisms, expanding the total parameter count of the Qianwen model to 2.4T for the first time, with 95B activated, achieving higher inference efficiency and faster inference speed, along with new breakthroughs in overall performance: In coding agent benchmarks such as PaperBench for scientific replication evaluation, Qwen3.8-Max improved by 28.2 points over the previous generation, setting a new benchmark high of 93.0 points; in general agent benchmarks like WideSearch and Agent's Last Exam, the new model scored 81.9 and 52.4 points, respectively, leading the frontier. Additionally, Qwen3.8 achieved 82.8 points in the instruction-following IF Bench evaluation and 92.6 points in the scientific reasoning GPQA Diamond, ranking among the top in a range of general capabilities; in the visual reasoning BabyVision evaluation, Qwen3.8-Max scored 82.0 points without tools, nearly double that of some mainstream models, and in the OSWorld-Verified benchmark assessing agent computer operation capabilities, Qwen3.8-Max ranked first among mainstream models with 86.1 points.
Coding capability is one of the core abilities of frontier large models, and Qwen3.8 has achieved a new breakthrough in autonomous programming. In the authoritative CodeArena ranking, Qwen3.8 ranks fourth globally. Two years ago, frontier models could write functions and complete code, serving as powerful assistants to programmers. Today, Qwen3.8 can start from an empty folder and autonomously deliver a real project that would take over ten days, without any human intervention throughout the process. With just a command to "create a self-evolving agent Harness," Qwen3.8 acts like a senior engineer, building a Loop Engineering framework from scratch, orchestrating agents to automatically claim and execute tasks, while human engineers can also submit new requirements to Qwen3.8 via DingTalk. After running autonomously for approximately 16 days, Qwen3.8 produced a real, usable self-evolving agent framework called "oh-my-cli" at the level of Hermes Agent. The project is now fully open-sourced, with the complete process publicly stored in a GitHub repository for anyone to review.
Qwen3.8 is capable of handling a wide range of real-world professional tasks (Cowork). Facing entirely different professional tasks, evaluation criteria, and computational demands, Qwen3.8 achieves improved real-world performance across hundreds of high-value, high-frequency professional tasks through joint reinforcement learning (RL) scaling with real environments and compute, delivering production-grade results reliably across mainstream agent frameworks: A legal assistant team annotating over a thousand relevant clauses across hundreds of legal documents would take a week; Qwen3.8 completes it within an hour. A basketball data analysis team annotating over 160 hours of game footage frame by frame; Qwen3.8 can precisely break down over 8,400 offensive and defensive plays in tens of minutes, outputting player tactical profiles and coach reports in one pass. A financial research team, relying on years of professional knowledge accumulation and comprehensive, real-time market data collection, would sequentially complete quantitative strategy research; Qwen3.8 autonomously orchestrates parallel agents, working continuously for several hours to deliver a complete ETF rotation strategy, continuously analyzing backtests, dynamically adjusting, and ultimately achieving scalable profitability.
In long-horizon tasks, Qwen3.8 exhibits emergent system-level autonomous planning and full-stack closed-loop adaptive learning capabilities. Whether in high-precision digital chip design with strong physical constraints or in highly dynamic, intensely competitive business simulation operations, Qwen3.8 leverages an adaptive closed loop of "execution-feedback-iteration" to achieve deep restructuring of algorithms and strategies across thousands of rounds of ultra-long-horizon interactions.
Beyond its powerful text and reasoning capabilities, Qwen3.8 also pushes the limits of AI visual understanding. In the authoritative Vision Arena ranking, Qwen3.8-Max ranks second globally. Qwen3.8 can not only read a 200-page financial report PDF and understand long videos exceeding 100 hours, but also feed this "seen" knowledge and information into the entire workflow, helping the model think more comprehensively. In the "application replication" benchmark RecreationBench for long-horizon tasks built by the Qianwen team, the model, without source code or internet access, understands an application only through interaction and feedback, yet can replicate the entire application from scratch. This demonstrates that Qwen3.8 has already exhibited frontier-level hybrid multimodal agent capabilities, pushing Visual Coding to new heights through repeated cycles of "iterative programming-interactive feedback."
It is understood that Alibaba's Zhenwu M890 supernode has also been successfully adapted for Qwen3.8, achieving significant improvements in inference efficiency and reductions in inference cost through full-stack optimization of chips, cloud platforms, and the Qianwen large model, with up to 1.5x performance gains in Agentic inference scenarios.









