the AI bench
VERIFIED AUGUST 2026
All models

MODEL · ALIBABA (QWEN) · 2.446T TOTAL / 95B ACTIVE PER TOKEN (SPARSE MOE; 92 LAYERS, 512 EXPERTS WITH 10 ROUTED + 1 SHARED, HIDDEN 8192)

Qwen3.8-Max (2.4T-A95B)

The first time Alibaba has opened a Max-tier model, and the resolution of a promise we said in July we would not take on faith. Qwen has a genuinely strong open-weight record, but the Max tier had never been opened — Qwen3.7-Max and 3.7-Plus both launched closed and stayed closed — so "Qwen3.8 will be open" was a reversal of strategy, not a continuation of it. The blog on August 3 said weights "next week"; they landed on the 8th. It is a 2.4-trillion-parameter sparse MoE with 95B active, built on the Qwen3.5 architecture with a Gated DeltaNet / Gated Attention hybrid layout, and Qwen positions it against Opus 4.8, Fable 5 and GPT-5.6 Sol. One distinction worth keeping straight: the open checkpoint is the text model, while the hosted Qwen3.8-Max adds vision input, non-thinking mode, 1M context by default and built-in tools. They are not the same artifact.

License: Custom "qwen3.8-max" licence (HF-tagged `other`) — MIT-like for most users, with two revenue clauses. NOT OSI-approved. · Context: 1M tokens on the hosted Qwen3.8-Max service; the open checkpoint is the post-trained text model · Released: Announced August 3, 2026; open weights published August 8, 2026 — the promise kept, one week later

The decision in five lines

The call
Hosted only
Best for
Hosted reference and benchmarks
Runs on
Hosted or workstation-class only · Not locally runnable — 2.45T parameters across 224 files
Watch out
Any local deployment, at any quantization — this is roughly the size of Kimi K3 and lives in the same datacenter-only bracket.
Evidence
Estimated · last verified August 2026

2.446T total
PARAMETERS
MOE
TYPE
1M
CONTEXT
Not locally runnable — 2.45T parameters across 224 files
VRAM AT Q4

Where we recommend this

This model isn’t currently in an active planner slot. See the runner notes below if you’re running it anyway.

The call

The first time Alibaba has opened a Max-tier model, and the resolution of a promise we said in July we would not take on faith. Qwen has a genuinely strong open-weight record, but the Max tier had never been opened — Qwen3.7-Max and 3.7-Plus both launched closed and stayed closed — so "Qwen3.8 will be open" was a reversal of strategy, not a continuation of it. The blog on August 3 said weights "next week"; they landed on the 8th. It is a 2.4-trillion-parameter sparse MoE with 95B active, built on the Qwen3.5 architecture with a Gated DeltaNet / Gated Attention hybrid layout, and Qwen positions it against Opus 4.8, Fable 5 and GPT-5.6 Sol. One distinction worth keeping straight: the open checkpoint is the text model, while the hosted Qwen3.8-Max adds vision input, non-thinking mode, 1M context by default and built-in tools. They are not the same artifact.

When not to use: Any local deployment, at any quantization — this is roughly the size of Kimi K3 and lives in the same datacenter-only bracket. Read the licence before commercial work, because it is not MIT despite being widely described that way: if your product exceeds 100M monthly active users or US$20M monthly revenue you must display the model name in your UI, and if you run a Model-as-a-Service or an AI coding/office-assistant business with over US$50M in any twelve months you need a separate licence from Qwen before using it commercially at all. For the overwhelming majority of readers neither clause binds, and the licence is effectively permissive — but "effectively permissive with revenue triggers" is not the same as OSI-open, and the difference matters the day it matters.

Runner notes

The BF16 repo is `Qwen/Qwen3.8-2.4T-A95B`; an FP8 variant ships alongside it and is the more downloaded of the two. Qwen states the artifacts are compatible with vLLM, SGLang and TokenSpeed. Reasoning depth is tunable via `reasoning_effort` (xhigh / medium / low) and `preserve_thinking` retains reasoning context across turns — it is on by default on the hosted service. Practically you consume this through Qwen Cloud (`qwen3.8-max`, OpenAI- and Anthropic-compatible endpoints), not on your own hardware. We list it as a frontier reference point, not a pick. The pattern worth noting: three of the largest open-weight releases of 2026 — Kimi K3, this, and DeepSeek V4-Pro — all landed within a month, and only DeepSeek's is under a classic OSI licence.

License
Custom "qwen3.8-max" licence (HF-tagged `other`) — MIT-like for most users, with two revenue clauses. NOT OSI-approved.
Released
Announced August 3, 2026; open weights published August 8, 2026 — the promise kept, one week later
Maker
Alibaba (Qwen)

Next step

Find-by-model — see what hardware runs this