MODEL · ALIBABA (QWEN) · 2.446T TOTAL / 95B ACTIVE PER TOKEN (SPARSE MOE; 92 LAYERS, 512 EXPERTS WITH 10 ROUTED + 1 SHARED, HIDDEN 8192)
Qwen3.8-Max (2.4T-A95B)
The first time Alibaba has opened a Max-tier model, and the resolution of a promise we said in July we would not take on faith. Qwen has a genuinely strong open-weight record, but the Max tier had never been opened — Qwen3.7-Max and 3.7-Plus both launched closed and stayed closed — so "Qwen3.8 will be open" was a reversal of strategy, not a continuation of it. The blog on August 3 said weights "next week"; they landed on the 8th. It is a 2.4-trillion-parameter sparse MoE with 95B active, built on the Qwen3.5 architecture with a Gated DeltaNet / Gated Attention hybrid layout, and Qwen positions it against Opus 4.8, Fable 5 and GPT-5.6 Sol. One distinction worth keeping straight: the open checkpoint is the text model, while the hosted Qwen3.8-Max adds vision input, non-thinking mode, 1M context by default and built-in tools. They are not the same artifact.
License: Custom "qwen3.8-max" licence (HF-tagged `other`) — MIT-like for most users, with two revenue clauses. NOT OSI-approved. · Context: 1M tokens on the hosted Qwen3.8-Max service; the open checkpoint is the post-trained text model · Released: Announced August 3, 2026; open weights published August 8, 2026 — the promise kept, one week later
The decision in five lines
- The call
- Hosted only
- Best for
- Hosted reference and benchmarks
- Runs on
- Hosted or workstation-class only · Not locally runnable — 2.45T parameters across 224 files
- Watch out
- Any local deployment, at any quantization — this is roughly the size of Kimi K3 and lives in the same datacenter-only bracket.
- Evidence
- Estimated
- 2.446T total
- PARAMETERS
- MOE
- TYPE
- 1M
- CONTEXT
- Not locally runnable — 2.45T parameters across 224 files
- VRAM AT Q4
Where we recommend this
This model isn’t currently in an active planner slot. See the runner notes below if you’re running it anyway.
The call
The first time Alibaba has opened a Max-tier model, and the resolution of a promise we said in July we would not take on faith. Qwen has a genuinely strong open-weight record, but the Max tier had never been opened — Qwen3.7-Max and 3.7-Plus both launched closed and stayed closed — so "Qwen3.8 will be open" was a reversal of strategy, not a continuation of it. The blog on August 3 said weights "next week"; they landed on the 8th. It is a 2.4-trillion-parameter sparse MoE with 95B active, built on the Qwen3.5 architecture with a Gated DeltaNet / Gated Attention hybrid layout, and Qwen positions it against Opus 4.8, Fable 5 and GPT-5.6 Sol. One distinction worth keeping straight: the open checkpoint is the text model, while the hosted Qwen3.8-Max adds vision input, non-thinking mode, 1M context by default and built-in tools. They are not the same artifact.
When not to use: Any local deployment, at any quantization — this is roughly the size of Kimi K3 and lives in the same datacenter-only bracket. Read the licence before commercial work, because it is not MIT despite being widely described that way: if your product exceeds 100M monthly active users or US$20M monthly revenue you must display the model name in your UI, and if you run a Model-as-a-Service or an AI coding/office-assistant business with over US$50M in any twelve months you need a separate licence from Qwen before using it commercially at all. For the overwhelming majority of readers neither clause binds, and the licence is effectively permissive — but "effectively permissive with revenue triggers" is not the same as OSI-open, and the difference matters the day it matters.
Runner notes
The BF16 repo is `Qwen/Qwen3.8-2.4T-A95B`; an FP8 variant ships alongside it and is the more downloaded of the two. Qwen states the artifacts are compatible with vLLM, SGLang and TokenSpeed. Reasoning depth is tunable via `reasoning_effort` (xhigh / medium / low) and `preserve_thinking` retains reasoning context across turns — it is on by default on the hosted service. Practically you consume this through Qwen Cloud (`qwen3.8-max`, OpenAI- and Anthropic-compatible endpoints), not on your own hardware. We list it as a frontier reference point, not a pick. The pattern worth noting: three of the largest open-weight releases of 2026 — Kimi K3, this, and DeepSeek V4-Pro — all landed within a month, and only DeepSeek's is under a classic OSI licence.
Next step
Find-by-model — see what hardware runs this→