the AI bench
VERIFIED AUGUST 2026
All fast takes

FAST TAKE · 2026-07-19 · ALIBABA QWEN3.8-MAX-PREVIEW (2.4T)

Qwen3.8-Max — a 2.4T claim with nothing published to check it against

Alibaba previewed Qwen3.8-Max on July 19 at WAIC Shanghai: 2.4 trillion parameters, its first multimodal model above 1T, and — unusually for the Max tier — a promise of open weights "soon". It is API-only today via QwenCloud's Token Plan, Qoder and QoderWork. The claim attached to it is that the model is "second only to Fable 5". There is no benchmark table, no model card, and no disclosed active-parameter count backing that. We are covering it because it is the biggest Qwen story of the year, and flagging plainly that almost nothing about it is yet checkable.

Verdict: Alibaba's biggest model ever and an open-weight promise — but no benchmark table, no model card, no active-parameter count, and no weights


The take

What is actually confirmed: the announcement came from the official @Alibaba_Qwen account on July 19, 2026, during WAIC in Shanghai — a post, not a paper or press release. The preview is real and purchasable, model ID `qwen3.8-max-preview`, listed in QwenCloud's Token Plan (Personal and Team) allowlist with visual understanding, text generation, thinking support, function calling and a 1M-context category. Qwen developer Shuai Bai confirmed sparse MoE and that this is Alibaba's first multimodal model above a trillion parameters. Preview access is offered at 10% of standard pricing, and there is no published per-token rate card — which is why it gets no row in our cost calculator.

What is claimed but unverified, and the distinction matters: the "second only to Fable 5" line is the entire public evidence base for the ranking. No scores, no named suite, no methodology, and no independent scoring from Artificial Analysis or LMArena at the time of writing. The 2.4T figure is vendor-reported with no model card behind it. Most consequentially for anyone thinking about deployment, **the active-parameter count is undisclosed** — and for a sparse MoE that is the single number that determines serving cost and latency. A 2.4T model with 40B active and one with 200B active are entirely different propositions; until Alibaba publishes it, any claim about what this costs to run is a guess.

On the open-weight promise, some earned scepticism. Alibaba has a genuinely strong open-weight record — Qwen is the most-used open family on our own site, and much of it ships Apache 2.0. But the **Max tier specifically has never been opened**: Qwen3.7-Max and Qwen3.7-Plus both launched closed earlier in 2026 and remain closed. So "Qwen3.8 will be open" is a reversal of Max-tier strategy, not a continuation of it, and no date, licence, or confirmation that the released checkpoint would be the same 2.4T Max checkpoint has been given. The confirming artifact is a Hugging Face repo with a licence file and real safetensors — exactly what we waited for with Kimi K3, which did deliver on July 27.

Our call: fast take, no model entry, no planner pick, no cost-calculator row. Even if the weights land, 2.4T at 4-bit is roughly 1.2 TB of weights before cache and runtime overhead — past the combined memory of eight top-tier accelerators, and an order of magnitude beyond every box on our hardware list. **The open-weight ceiling for the Qwen line remains Qwen 3.6** (35B-A3B and 27B, Apache 2.0); we re-checked the Hugging Face API on August 3 and the newest Qwen repo of any kind is still 2026-06-26. The pattern worth noticing is the fortnight itself: Kimi K3 at 2.8T on July 16, Qwen3.8 at 2.4T on July 19, each claiming a spot just below the closed frontier. Two of those three shipped weights. We will update this take the day Qwen ships theirs, or the day the promise quietly expires.

Where this fits

Models: Qwen 3.6-35B-A3B · Qwen 3.6-27B · Kimi K3 (2.8T-A50B) · Qwen 3.5 122B-A10B

Hardware: NVIDIA DGX Spark · Dual RTX 5090 · Mac Studio M3 Ultra 96 GB

Sources

Next step

Try this in the planner