AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: Qwen3.8-Max's AI Data: A Step Closer To Top Spot Or Just A Fluke? on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba announced the broad availability of Qwen3.8-Max, a 2.4-trillion-parameter AI model, with confirmed benchmark results. Its performance is impressive but selective, raising questions about its true standing in the AI landscape.

Alibaba has officially released the full benchmark data for its AI model Qwen3.8-Max, confirming it features 2.4 trillion parameters and is now broadly available for testing and deployment. This development marks a significant step in the AI industry, as the model demonstrates competitive performance in several key benchmarks, positioning it as a potential contender for the top spot among large language models.

On August 3, Alibaba confirmed that Qwen3.8-Max, previously only previewed in stealth, now has its full benchmark table published and open weights scheduled for next week. The model boasts 2.4 trillion total parameters, with approximately 95 billion active parameters per query, built on a sparse mixture-of-experts architecture based on Qwen3.5. It is multimodal, supporting text, image, and video inputs, with text output.

The benchmark results reveal a model that performs strongly in several areas. It scores 86.6 on Terminal-Bench 2.1, surpassing Claude models and only behind GPT-5.6 Sol at 88.8, and achieves top marks on PaperBench at 93.0. It also excels in multimodal and agentic tasks, with notable scores on OSWorld-Verified and Parametric CAD Bench. However, it trails significantly on deep software engineering benchmarks like SWE-bench Pro and FrontierSWE, with gaps of 12-15 points compared to Fable 5.

The open weights are considered more of a gesture than an immediate deployment option due to their size, requiring multi-node datacenter infrastructure. A smaller, 27B checkpoint, Qwen3.8-27B, is also available, optimized for single high-memory machines, and is expected to perform well on inference tasks relevant to end-users.

At a glance
updateWhen: announced August 3, 2023; benchmarks an…
The developmentAlibaba made Qwen3.8-Max publicly accessible with full benchmark data, confirming it as one of the largest open-weight models with strong performance in key tests.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba's Open-Weight Release

The release of Qwen3.8-Max's full benchmark data and upcoming open weights signals a major move in AI openness and competition. While the model demonstrates impressive performance in several benchmarks, its selective strengths suggest it is not yet a definitive top-tier model across all tasks. The availability of the 27B checkpoint for local deployment could influence real-world AI applications, especially if agentic and long-horizon capabilities are preserved in compression.

This development intensifies the race among large AI firms to deliver models that are both powerful and accessible, with Alibaba positioning itself as a serious contender. However, the true impact depends on whether the model's strengths translate into broad utility outside benchmark settings and how licensing and deployment restrictions evolve.

LLM Systems Engineering: Training and Building Large Language Models – Engineering AI Models Through Fine-Tuning, Continued Pretraining, and From-Scratch Development

LLM Systems Engineering: Training and Building Large Language Models – Engineering AI Models Through Fine-Tuning, Continued Pretraining, and From-Scratch Development

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of Alibaba's AI Model Launches

Over the past two weeks, Alibaba's AI developments have unfolded through a series of strategic moves. On July 17, the company previewed Kimi K3, a 2.8 trillion-parameter model that briefly affected US tech stocks due to high demand. The following day, an anonymous model called 'kaleb' appeared on the Code Arena leaderboard, later confirmed as Qwen3.8-Max during the World AI Conference in Shanghai on July 19. This stealth reveal was accompanied by claims of being 'second only to Fable 5,' though without concrete benchmarks until now.

Until August 3, Alibaba maintained a low profile, releasing limited preview endpoints and withholding full benchmark data. The recent publication of the benchmark table and the scheduled release of open weights mark a shift from stealth to transparency, allowing industry experts to evaluate the model's true capabilities and competitive standing.

"We are committed to advancing AI openness and providing developers with powerful tools, starting with the broad availability of Qwen3.8-Max's weights next week."

— Alibaba spokesperson

ASRock Intel Arc Pro B70 Creator 32GB Workstation Graphics Card, Xe2-HPG, 32GB GDDR6, PCIe 5.0, 4X DP 2.1, Blower Fan, Vapor Chamber, Honeywell PTM7950

ASRock Intel Arc Pro B70 Creator 32GB Workstation Graphics Card, Xe2-HPG, 32GB GDDR6, PCIe 5.0, 4X DP 2.1, Blower Fan, Vapor Chamber, Honeywell PTM7950

  • System Compatibility: Measures 271 x 112 x 39 mm
  • Power Requirements: Requires 12V-2x6-pin connector
  • Customer Support: Direct Amazon contact for assistance

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About Model Performance

It is not yet clear whether the high benchmark scores, especially in multimodal and agentic tasks, will translate into real-world applications. The impact of model compression on agentic capabilities remains uncertain, as does the final licensing framework for open weights, which could influence deployment options.

Additionally, the full scope of the model's performance on tasks outside the benchmark suite, and how it compares to other leading models in practical settings, is still to be seen.

Building Intelligent Applications with Spring AI: Develop Practical Java Solutions with Generative AI, Multimodal Models, and Agents

Building Intelligent Applications with Spring AI: Develop Practical Java Solutions with Generative AI, Multimodal Models, and Agents

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Upcoming Steps for Alibaba’s AI Strategy

Next week, Alibaba plans to release the open weights for Qwen3.8-27B, which is expected to be more accessible for local deployment. Industry observers will evaluate whether the smaller checkpoint retains the agentic and long-horizon capabilities demonstrated by the full model.

Further benchmarking and real-world testing will clarify the model’s true standing and influence in the AI ecosystem. Additionally, the licensing terms and potential integration into commercial products will shape its broader adoption.

RackChoice 3U rackmount Server Chassis Support Liquid Cooling Compatibility up to Elevated 360mm Radiator Support SFX PSU/ATX/MicroATX/Mini-ITX MB

RackChoice 3U rackmount Server Chassis Support Liquid Cooling Compatibility up to Elevated 360mm Radiator Support SFX PSU/ATX/MicroATX/Mini-ITX MB

  • Pre-installed 120mm Fans: Includes 3 fans or supports 360mm radiator
  • Motherboard Compatibility: Supports ATX, MicroATX, Mini-ITX
  • Drive Bays: 2×3.5-inch and 1×2.5-inch internal bays

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the significance of Alibaba releasing Qwen3.8-Max's benchmark data?

The release provides transparency on the model's capabilities and allows industry experts to evaluate its performance across various benchmarks, helping determine its competitiveness in the AI landscape.

Will the open weights allow broader access for developers?

Yes, Alibaba announced that the open weights for Qwen3.8-27B will be available next week, enabling local deployment on high-memory machines, though the full 2.4 trillion-parameter model remains a datacenter artifact.

Does the model outperform competitors across all tasks?

No, while it excels in certain benchmarks, especially multimodal and agentic tasks, it trails significantly in deep software engineering benchmarks, indicating selective strengths rather than universal dominance.

What are the licensing implications for the open weights?

The licensing terms are still unpublished. Historically, Alibaba's open models have used Apache 2.0, but the upcoming license could include revenue or attribution triggers, affecting deployment and commercial use.

How might this development influence the AI industry?

It could accelerate competition by setting a new benchmark for openness and performance, prompting other firms to release larger models or improve transparency, ultimately benefiting AI developers and users.

Source: ThorstenMeyerAI.com

You May Also Like

AI Trading Bot — Week Two: The candidate edge collapsed

The promising BTC fair-value strategy lost nearly all its gains in week two, with all tested approaches now in the red, highlighting the fragility of AI trading edges.

Own Your AI Model Like A Pro With Tinker, Forge, Or Frontier Tuning

Discover how Tinker, Forge, and Frontier Tuning enable organizations to control their AI models securely and compliantly, tailored for regulated sectors.

The Mechanics Behind Claude Watermark: A New AI Text-Marking Method

A new report suggests Anthropic’s Claude may use a novel text-marking method, but technical details and deployment status remain unconfirmed.

The Stanford AI Index 2026 Audit: Reading the Field’s Annual Report Card With a Critic’s Pen

The Stanford AI Index 2026 has been released, offering a comprehensive but partial snapshot of AI progress. This analysis examines its strengths, limitations, and implications.