the AI bench
VERIFIED AUGUST 2026
All models

MODEL · ALIBABA · 27.8B DENSE

Qwen3.8-27B

The August 2026 dense refresh, and the most-adopted open-weight model of the month by a wide margin — 267K downloads and 10.4k likes in its first twelve days, with 480 community quantizations already published. It supersedes Qwen 3.6-27B at the same size, the same Apache 2.0 licence and the same 24 GB target, but the architecture underneath is different: a hybrid layout of 16 × (3 × Gated DeltaNet → 1 × Gated Attention), the same linear-attention family Alibaba used in Qwen3.8-Max. Three quarters of its 64 layers use linear attention whose state does not grow with context, so the KV cache at long context is roughly half what the dense 3.6-27B needs. It is also a native vision-language model — images and video in, not a text model with an adapter bolted on — and it is trained with multi-token prediction, so speculative decoding works without a separate drafter.

License: Apache 2.0 · Context: 262K native, extensible to 1M · Released: August 5, 2026

The decision in five lines

The call
Buy — for chat
Best for
chat · docs
Runs on
16 hardware picks fit (cheapest: Minisforum UM890 Pro · $463)
Watch out
When you want the fastest tokens/sec on a 24 GB card — the 3.6-35B-A3B MoE still wins on raw decode speed because only 3B parameters activate per token.
Evidence
Estimated · last verified August 2026

27.8B dense
PARAMETERS
DENSE
TYPE
262K
CONTEXT
~17 GB
VRAM AT Q4

Where we recommend this

Every tier slot in the planner where this model is a top or alternate pick. Pulled live from planner.js — when the planner refreshes, this table stays current.

CHAT · TOP
Qwen3.8-27BAugust 5 2026 dense refresh, Apache 2.0, ~17 GB at Q4. Native vision, and a Gated DeltaNet hybrid that keeps a growing KV cache on only 16 of 64 layers — roughly half the cache of the 3.6-27B it replaces.
DOCS · TOP
Qwen3.8-27B262K native context extensible to 1M, native vision, single-GPU at Q4. The linear-attention hybrid matters most here: three quarters of its layers hold a constant-size state, so long-doc work costs far less memory than a conventional dense 27B.

The call

The August 2026 dense refresh, and the most-adopted open-weight model of the month by a wide margin — 267K downloads and 10.4k likes in its first twelve days, with 480 community quantizations already published. It supersedes Qwen 3.6-27B at the same size, the same Apache 2.0 licence and the same 24 GB target, but the architecture underneath is different: a hybrid layout of 16 × (3 × Gated DeltaNet → 1 × Gated Attention), the same linear-attention family Alibaba used in Qwen3.8-Max. Three quarters of its 64 layers use linear attention whose state does not grow with context, so the KV cache at long context is roughly half what the dense 3.6-27B needs. It is also a native vision-language model — images and video in, not a text model with an adapter bolted on — and it is trained with multi-token prediction, so speculative decoding works without a separate drafter.

When not to use: When you want the fastest tokens/sec on a 24 GB card — the 3.6-35B-A3B MoE still wins on raw decode speed because only 3B parameters activate per token. Also treat the benchmark tables on the model card with the usual caution: they are vendor-run, evaluated with Alibaba's own harness, and we have not seen independent numbers yet. The reason this displaces Qwen 3.6-27B in our picks is not the benchmark claims — it is that it is the same vendor's direct successor at the same size and licence, with an order of magnitude more adoption and a materially cheaper long-context profile.

Runner notes

Native Ollama tag `qwen3.8:27b` already resolves (verified 2026-08-17) — unusually fast for a Qwen release, which historically took 1–2 weeks. `unsloth/Qwen3.8-27B-GGUF` is the community quant of record at ~1.9M downloads; MLX builds exist for Apple Silicon. Use vLLM or SGLang for long-context serving. To go past 262K, set the `factor` in `rope_parameters` under `text_config` in `config.json` — Alibaba suggests matching it to your typical context, e.g. factor 2.0 for ~524K. The wide quant mirroring matters for more than convenience: after the Microsoft Lens withdrawal in August, a model living in hundreds of community repos is far harder to make disappear than one that lives only in a vendor org.

License
Apache 2.0
Released
August 5, 2026
Maker
Alibaba

Hardware that fits

Every hardware pick whose memory fits this model at the quant we recommend. Sorted cheapest-first — the top row is your best-value fit. Click through for the full buyer’s guide.

Next step

Find-by-model — see what hardware runs this