MODEL · ALIBABA · 27.8B DENSE
Qwen3.8-27B
The August 2026 dense refresh, and the most-adopted open-weight model of the month by a wide margin — 267K downloads and 10.4k likes in its first twelve days, with 480 community quantizations already published. It supersedes Qwen 3.6-27B at the same size, the same Apache 2.0 licence and the same 24 GB target, but the architecture underneath is different: a hybrid layout of 16 × (3 × Gated DeltaNet → 1 × Gated Attention), the same linear-attention family Alibaba used in Qwen3.8-Max. Three quarters of its 64 layers use linear attention whose state does not grow with context, so the KV cache at long context is roughly half what the dense 3.6-27B needs. It is also a native vision-language model — images and video in, not a text model with an adapter bolted on — and it is trained with multi-token prediction, so speculative decoding works without a separate drafter.
License: Apache 2.0 · Context: 262K native, extensible to 1M · Released: August 5, 2026
The decision in five lines
- The call
- Buy — for chat
- Best for
- chat · docs
- Runs on
- 16 hardware picks fit (cheapest: Minisforum UM890 Pro · $463)
- Watch out
- When you want the fastest tokens/sec on a 24 GB card — the 3.6-35B-A3B MoE still wins on raw decode speed because only 3B parameters activate per token.
- Evidence
- Estimated
- 27.8B dense
- PARAMETERS
- DENSE
- TYPE
- 262K
- CONTEXT
- ~17 GB
- VRAM AT Q4
Where we recommend this
Every tier slot in the planner where this model is a top or alternate pick. Pulled live from planner.js — when the planner refreshes, this table stays current.
The call
The August 2026 dense refresh, and the most-adopted open-weight model of the month by a wide margin — 267K downloads and 10.4k likes in its first twelve days, with 480 community quantizations already published. It supersedes Qwen 3.6-27B at the same size, the same Apache 2.0 licence and the same 24 GB target, but the architecture underneath is different: a hybrid layout of 16 × (3 × Gated DeltaNet → 1 × Gated Attention), the same linear-attention family Alibaba used in Qwen3.8-Max. Three quarters of its 64 layers use linear attention whose state does not grow with context, so the KV cache at long context is roughly half what the dense 3.6-27B needs. It is also a native vision-language model — images and video in, not a text model with an adapter bolted on — and it is trained with multi-token prediction, so speculative decoding works without a separate drafter.
When not to use: When you want the fastest tokens/sec on a 24 GB card — the 3.6-35B-A3B MoE still wins on raw decode speed because only 3B parameters activate per token. Also treat the benchmark tables on the model card with the usual caution: they are vendor-run, evaluated with Alibaba's own harness, and we have not seen independent numbers yet. The reason this displaces Qwen 3.6-27B in our picks is not the benchmark claims — it is that it is the same vendor's direct successor at the same size and licence, with an order of magnitude more adoption and a materially cheaper long-context profile.
Runner notes
Native Ollama tag `qwen3.8:27b` already resolves (verified 2026-08-17) — unusually fast for a Qwen release, which historically took 1–2 weeks. `unsloth/Qwen3.8-27B-GGUF` is the community quant of record at ~1.9M downloads; MLX builds exist for Apple Silicon. Use vLLM or SGLang for long-context serving. To go past 262K, set the `factor` in `rope_parameters` under `text_config` in `config.json` — Alibaba suggests matching it to your typical context, e.g. factor 2.0 for ~524K. The wide quant mirroring matters for more than convenience: after the Microsoft Lens withdrawal in August, a model living in hundreds of community repos is far harder to make disappear than one that lives only in a vendor org.
Hardware that fits
Every hardware pick whose memory fits this model at the quant we recommend. Sorted cheapest-first — the top row is your best-value fit. Click through for the full buyer’s guide.
- Minisforum UM890 ProGood · 1.3× 32 GB DDR5 (shared) · $463–$580 all-in
- AMD Radeon RX 7900 XTXGood · 1.3× 24 GB · $850 refurb / $1,400–$1,500 new
- NVIDIA RTX 3090 (used, single)Good · 1.3× 24 GB · $950–$1,200
- MacBook Air M5 24 GBRequires tweak · 1.2× 24 GB unified · $1,499–$1,899
- Mac Mini M4 Pro 24 GBRequires tweak · 1.2× 24 GB unified · $1,599
- Dual RTX 3090 (used)Perfect · 2.6× 48 GB · $1,800–$2,500 all-in
- NVIDIA RTX 4090Good · 1.3× 24 GB · $2,200–$2,800
- M5 Pro MacBook Pro 48 GBPerfect · 1.7× 48 GB unified · $2,999–$3,599
- Framework Desktop (Ryzen AI Max+ 395)Perfect · 4.6× 128 GB unified · $3,449 (128 GB config)
- NVIDIA RTX A6000 (48 GB, used)Perfect · 2.6× 48 GB ECC · $3,500–$4,500
- Mac Studio M4 Max 64 GBPerfect · 2.3× 64 GB unified · $3,799
- NVIDIA RTX 5090Perfect · 1.7× 32 GB · $4,000–$4,900
- NVIDIA DGX SparkPerfect · 4.6× 128 GB unified · $4,699
- M5 Max MacBook Pro 64 GBPerfect · 2.3× 64 GB unified · ~$5,199 (est.; June 25 2026 increase)
- Mac Studio M3 Ultra 96 GBPerfect · 3.5× 96 GB unified · $5,299
- Dual RTX 5090Perfect · 3.5× 64 GB (2×32) · $8,500–$10,500
Next step
Find-by-model — see what hardware runs this→