📊 Full opportunity report: Can Qwen3.8-Max Be The Second Best? The Data Offers Clues on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Alibaba announced the broad availability of Qwen3.8-Max, a 2.4 trillion-parameter model with strong benchmark performance. While it claims to be second only to Fable 5, the data shows it excels in some areas but trails in others, especially deep software benchmarks. The open weights will be released next week, with implications for AI deployment and open model development.
Alibaba has officially published benchmark results for Qwen3.8-Max, confirming it as a high-performing AI model with 2.4 trillion parameters. This marks the first time the model’s detailed performance data has been made public, and the open weights are set to ship next week, making it a significant development in the AI model landscape.
On August 3, Alibaba released the full benchmark table for Qwen3.8-Max, revealing a model built on the Qwen3.5 architecture with approximately 95 billion active parameters per query. The model employs sparse mixture-of-experts technology and supports multimodal inputs—text, images, and video—with text output. It achieved top scores on several benchmark tests, including Terminal-Bench 2.1 (86.6), PaperBench (93.0), and IFBench (82.8). Notably, it outperformed some competitors in multimodal and agentic tasks, such as OSWorld-Verified (86.1) and Parametric CAD Bench (91.5). However, it trails significantly in deep software engineering benchmarks like SWE-bench Pro (67.7) and FrontierSWE (73.5), compared to Fable 5’s higher scores.
The model’s performance indicates a roughly 95-billion-parameter active core within a 2.4-trillion-parameter network, utilizing sparse mixture-of-experts architecture. Alibaba also demonstrated its ability to reproduce research paper results and improve agentic tasks through reinforcement learning environment scaling. The company’s announcement included confirmation that the open weights—set for next week—will be accessible, though the licensing terms remain unpublished, raising questions about deployment options.
For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.
▲ All performance figures: Alibaba’s own harnessThe claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.
“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.
“Qwen3.8 is going open-weight” describes three things with very different deployment realities.
OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.
A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.
The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.
Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.
- The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
- More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
- If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
- The 27B sibling could become the best local agent model on hardware people already own.
- Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
- The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
- “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
- Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
and it says “second only” depends entirely on which row you read.
Implications of Alibaba's Benchmark Results and Open Release
The release of Qwen3.8-Max benchmark data and upcoming open weights mark a pivotal moment in AI development. The model's competitive scores, especially in multimodal and agentic tasks, suggest it could influence AI research and deployment. The openness of the weights allows broader access for developers and researchers, potentially accelerating innovation. However, its limitations in deep software engineering benchmarks highlight ongoing challenges in AI performance consistency. This development underscores the growing importance of transparency, scalability, and open models in the AI ecosystem.

AI Value Creators: Beyond the Generative AI User Mindset
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background and Recent AI Model Launches
Alibaba's Qwen series has been evolving rapidly, with the initial preview of Qwen3.8-Max appearing stealthily on July 19 during the World AI Conference in Shanghai. The model's announcement followed a series of high-profile releases, including Moonshot's Kimi K3 with 2.8 trillion parameters and the anonymous 'kaleb' model, later identified as Qwen3.8-Max. Alibaba's strategy involved a staged reveal, culminating in the full benchmark disclosure and confirmation of open weights. The company's approach contrasts with other major AI labs that often release models with limited transparency, emphasizing Alibaba's focus on openness and competitive benchmarking.
Historically, Alibaba's open models have shipped under Apache 2.0 licenses, but the licensing details for Qwen3.8-Max remain unpublished, raising questions about future deployment and commercial use. The model's architecture, based on sparse mixture-of-experts, is designed to optimize performance across multimodal and agentic tasks, aligning with recent trends in large-scale AI models.
"We are committed to transparency and open access, and the upcoming release of open weights will enable broader innovation in the AI community."
— Alibaba spokesperson
Unconfirmed Licensing Terms and Deployment Scope
The licensing details for Qwen3.8-Max remain unpublished, creating uncertainty about how the open weights can be used commercially or in proprietary settings. It is also unclear whether the open weights will include the full 2.4 trillion parameters or just the 95 billion active core, and how the model's performance will translate when compressed for single-machine deployment.
Further, the benchmark results, while comprehensive, do not cover all possible application scenarios, and the long-term stability and support for the open weights are still unknown.
Next Steps for Open Weights and Model Adoption
Alibaba plans to release the Qwen3.8-Max open weights next week, which will enable researchers and developers to evaluate its performance firsthand. The community will closely monitor how well the compressed 27B version performs in real-world applications and whether the agentic improvements are maintained. Additionally, the licensing terms will likely be clarified, influencing how the model can be integrated into commercial products. Further benchmark testing and comparative analyses are expected as the model gains broader access.
Key Questions
When will Alibaba release the open weights for Qwen3.8-Max?
The open weights are scheduled to be released next week, with Alibaba confirming the upcoming availability.
How does Qwen3.8-Max compare to other leading models?
Benchmark data shows Qwen3.8-Max is top of the table in some tests like PaperBench and performs well in multimodal and agentic tasks, but it trails behind Fable 5 in deep software engineering benchmarks.
What are the licensing implications of the open weights?
Licensing details remain unpublished, raising questions about usage rights, commercial deployment, and whether the model's full 2.4 trillion parameters will be available for open use.
Will the model's agentic capabilities be maintained in compressed versions?
This remains uncertain. The 27B checkpoint, designed for single-machine deployment, will be tested to see if it retains the agentic improvements demonstrated by the full model.
What impact could this have on AI research and industry?
The open release of a high-performance, large-scale model like Qwen3.8-Max could accelerate innovation, democratize access, and challenge existing market leaders, depending on licensing and deployment outcomes.
Source: ThorstenMeyerAI.com