the AI bench
VERIFIED SEPTEMBER 2026
All models

MODEL · ALIBABA · 27B DENSE

Qwen 3.6-27B

The April 2026 dense refresh that supersedes Qwen 3.5 27B — claims to beat the prior 397B MoE flagship on coding benchmarks while staying single-GPU deployable at Q4. Superseded in turn by Qwen3.8-27B (August 5, 2026), which holds the same size, licence and 24 GB target but switches to a Gated DeltaNet hybrid that roughly halves the KV cache at long context. This entry is kept as the dated record and as the fallback if you want plain full attention on every layer; new deployments should start from 3.8.

License: Apache 2.0 · Context: 262K native, extendable to ~1M via YaRN · Released: April 22, 2026

The decision in five lines

The call
Consider — runnable locally, family reference
Best for
Local evaluation and family reference
Runs on
16 hardware picks fit (cheapest: Minisforum UM890 Pro · $463 barebone (no RAM / SSD); $991 with 32 GB / 1 TB — Minisforum US store, Sep 15 2026)
Watch out
When you need fastest possible tok/s on 24 GB VRAM — the 3.6-35B-A3B MoE sibling runs ~6× faster for comparable quality because only 3B activate per token.
Evidence
Estimated · last verified September 2026

27B dense
PARAMETERS
DENSE
TYPE
262K
CONTEXT
~17 GB
VRAM AT Q4

Where we recommend this

This model isn’t currently in an active planner slot. See the runner notes below if you’re running it anyway.

The call

The April 2026 dense refresh that supersedes Qwen 3.5 27B — claims to beat the prior 397B MoE flagship on coding benchmarks while staying single-GPU deployable at Q4. Superseded in turn by Qwen3.8-27B (August 5, 2026), which holds the same size, licence and 24 GB target but switches to a Gated DeltaNet hybrid that roughly halves the KV cache at long context. This entry is kept as the dated record and as the fallback if you want plain full attention on every layer; new deployments should start from 3.8.

When not to use: When you need fastest possible tok/s on 24 GB VRAM — the 3.6-35B-A3B MoE sibling runs ~6× faster for comparable quality because only 3B activate per token. The dense 27B trades speed for the simpler dense-attention behavior some workflows prefer.

Runner notes

GGUFs available via unsloth and bartowski. Native Ollama tag may take 1–2 weeks to land. Use vLLM or llama.cpp for long-context work; MLX-community builds for Apple Silicon.

License
Apache 2.0
Released
April 22, 2026
Maker
Alibaba

Hardware that fits

Every hardware pick whose memory fits this model at the quant we recommend. Sorted cheapest-first — the top row is your best-value fit. Click through for the full buyer’s guide.

Next step

Find-by-model — see what hardware runs this