📊 Full opportunity report: Breaking Down Qwen3.8-Max’s AI Capabilities: The Numbers You Need To Know on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba officially announced Qwen3.8-Max, revealing its specifications and benchmark results. The model features 2.4 trillion parameters with 95 billion active per query and outperforms many competitors in key benchmarks. Open weights will be available next week, enabling wider access.

Alibaba has officially confirmed the specifications and benchmark performance of Qwen3.8-Max, its largest-ever AI model, making it broadly available and revealing detailed metrics that were previously withheld. For more context, see Breaking Down The 19-Day Timeline Of AI Gate Closures.

On August 3, Alibaba published the full benchmark table for Qwen3.8-Max, confirming it has 2.4 trillion total parameters with approximately 95 billion active parameters per query. The model is based on a sparse mixture-of-experts architecture derived from Qwen3.5, and it supports multimodal input — text, images, and video — with text output. The active-parameter figure clarifies that the model’s core computation involves about 4% of the total parameters firing at once.

Benchmark results show strong performance: Terminal-Bench 2.1 at 86.6, outperforming Claude Opus 4.8 and Claude Fable 5, but slightly below GPT-5.6 Sol at 88.8. You can learn more about AI benchmarks and their significance in our Breaking Down The 19-Day Timeline Of AI Gate Closures article. In addition, PaperBench scores 93.0, placing it at the top of the table, while multimodal and agentic tasks highlight its capabilities, with scores like 86.1 on OSWorld-Verified and 91.5 on Parametric CAD Bench. The model also demonstrated advanced reasoning by reproducing research results and outperforming its predecessor in long-horizon agent tasks.

Alibaba announced that the open weights for the model will be released next week, although the current deployment remains a cloud-based API compatible with OpenAI and DashScope. The 2.4 trillion-parameter checkpoint is a multi-node data center artifact, making self-hosting impractical for most users. Meanwhile, a 27B version, Qwen3.8-27B, designed for local deployment on high-memory machines, will also be available, though benchmark data for this smaller model has not yet been published. For related insights, see Breaking Down The 19-Day Timeline Of AI Gate Closures.

At a glance
updateWhen: announced August 3, 2023; benchmarks pu…
The developmentAlibaba confirmed the broad availability of Qwen3.8-Max, publishing detailed benchmarks and announcing open weights release scheduled for next week.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba’s New AI Model Release

The announcement confirms Alibaba’s position as a major player in large-scale AI, with Qwen3.8-Max setting new benchmarks in multimodal and agentic tasks. The detailed specs and benchmark scores demonstrate significant advancements in AI capabilities, especially in long-horizon reasoning and multimodal understanding, which could influence future AI development and deployment strategies. The upcoming open weights release will enable broader access, potentially accelerating innovation and competition in the AI ecosystem.

However, the fact that the full model remains a multi-node data center artifact underscores ongoing challenges in democratizing access to such large models. The smaller 27B variant aims to bridge this gap, but its performance benchmarks are still awaited. The model’s agentic improvements, driven by reinforcement learning environment scaling, mark a notable step forward, though questions about the longevity of these gains through compression remain.

Amazon

AI model hosting hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Alibaba’s AI Model Development

Alibaba’s AI journey has been marked by strategic releases and stealth previews, culminating in the recent unveiling of Qwen3.8-Max. Prior to this, the company hinted at a model larger than 2.8 trillion parameters via a slogan and a brief preview, but detailed specifications and benchmarks were withheld until now. The model was initially identified as kaleb during a community-driven discovery, and Alibaba confirmed its identity during the World AI Conference in Shanghai on July 19.

The model’s development builds on Alibaba’s previous architectures, notably Qwen3.5, and incorporates sparse mixture-of-experts to manage its immense size. The launch strategy involved a staged reveal, including a paid preview endpoint and selective benchmarking, which built anticipation and speculation. The release of detailed benchmarks and the upcoming open weights mark a significant milestone in Alibaba’s AI roadmap, aligning with industry trends toward transparency and open access for large models.

"We are committed to advancing AI capabilities and providing broader access through open weights, enabling innovation across the industry."

— Alibaba spokesperson

Amazon

high-memory GPU for AI training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About Model Deployment and Performance

While benchmark scores and specifications are now confirmed, the performance of Qwen3.8-27B in real-world applications remains unreported. The licensing terms for the open weights are still unpublished, raising questions about usage rights and restrictions. Additionally, the durability of agentic gains after model compression has yet to be verified, leaving some uncertainty about practical deployment and long-term performance.

Amazon

AI development workstation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Upcoming Release and Industry Impact of Open Weights

Next week, Alibaba will release the open weights for Qwen3.8-Max, enabling wider access and testing by third-party developers. Industry analysts will closely monitor how the model performs across diverse tasks and whether the agentic improvements are sustained in smaller versions like Qwen3.8-27B. Further benchmark publications and licensing details are expected in the coming weeks, which will clarify the model’s potential for broader adoption and influence on AI development trends.

Amazon

multimodal AI input devices

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What are the key specifications of Qwen3.8-Max?

Qwen3.8-Max has 2.4 trillion total parameters, with about 95 billion active per query, built on a sparse mixture-of-experts architecture. It supports multimodal input and has demonstrated strong benchmark performance.

When will the open weights for Qwen3.8-Max be available?

The open weights are scheduled to be released next week, enabling broader access for research and deployment.

How does Qwen3.8-Max compare to other models in benchmarks?

It outperforms Claude models and is competitive with GPT-5.6 Sol in certain benchmarks, especially in multimodal and agentic tasks, though it trails behind the very top models in some software engineering benchmarks.

What are the implications for developers and industry?

The release of open weights will facilitate wider experimentation and deployment, potentially accelerating AI innovation, but practical limitations remain due to the model’s size and infrastructure requirements.

What uncertainties remain about Qwen3.8-Max?

Unanswered questions include the licensing terms, actual performance of the smaller Qwen3.8-27B model, and whether agentic gains will persist after model compression and deployment in real-world settings.

Source: ThorstenMeyerAI.com

You May Also Like

13 Best Guides to AI-Powered Marketing Automation Tools for Smarter Campaigns in 2026

Explore the 13 best guides for AI-driven marketing automation, helping businesses optimize campaigns, workflows, and growth strategies effectively.

GPT-5.6 Sol, Along With Terra And Luna, Will Launch Publicly This Thursday

GPT-5.6 Sol, along with Terra and Luna, will be publicly launched this Thursday, marking a significant event in AI and blockchain integration.

2X, Not 10X: Coding With LLMs In 2026

New analysis suggests language models in 2026 offer only 2x improvements in coding tasks, challenging expectations of 10x gains.

AI in Entertainment: Hollywood Strikes a Balance Between Tech and Talent

Probing how AI reshapes Hollywood reveals the delicate balance between technological innovation and talent integrity, prompting questions about the future of entertainment.