MODEL · DEEPSEEK · 284B TOTAL / 13B ACTIVE (MOE); THE 0731 REFRESH SHIPS AT 304B TOTAL
DeepSeek V4-Flash
The smaller half of the V4 family — 284B MoE with 13B active per token. Same 1M context, same MIT license, same architectural KV-cache improvements as V4-Pro. The honest local pick of the V4 line: still frontier-class on most benchmarks, but realistically deployable only on M3 Ultra 192GB unified or dual 80GB server cards.
License: MIT · Context: 1M tokens (384K max output) · Released: April 24, 2026 (preview); refreshed as DeepSeek-V4-Flash-0731 on July 31, 2026
The decision in five lines
- The call
- Hosted only
- Best for
- coding
- Runs on
- Hosted or workstation-class only · ~158 GB (Unsloth Q4_K_M; needs M3 Ultra 192GB+ unified or dual 80GB server cards; not single-card consumer)
- Watch out
- At ~158 GB Q4 this exceeds every pick in the planner's 22-card library — workstation tier, not consumer.
- Evidence
- Estimated
- 284B total
- PARAMETERS
- MOE
- TYPE
- 1M
- CONTEXT
- ~158 GB (Unsloth Q4_K_M; needs M3 Ultra 192GB+ unified or dual 80GB server cards; not single-card consumer)
- VRAM AT Q4
Where we recommend this
Every tier slot in the planner where this model is a top or alternate pick. Pulled live from planner.js — when the planner refreshes, this table stays current.
The call
The smaller half of the V4 family — 284B MoE with 13B active per token. Same 1M context, same MIT license, same architectural KV-cache improvements as V4-Pro. The honest local pick of the V4 line: still frontier-class on most benchmarks, but realistically deployable only on M3 Ultra 192GB unified or dual 80GB server cards.
When not to use: Single-card consumer hardware. At ~158 GB Q4 this exceeds every pick in the planner's 22-card library — workstation tier, not consumer. Use V4-Pro via API for outright frontier; use Qwen3-Coder-30B-A3B locally if you need single-card.
Runner notes
**Successor — September 10, 2026:** `deepseek-ai/DeepSeek-V4.1-Flash` (MIT, 552B backbone with 8B active for input and 16B for output, native image input, 1M context; 763B in safetensors because it counts a 196B "Engram" memory table) replaced this model on DeepSeek's API the same day. The API name is now `deepseek-flash`; `deepseek-v4-flash` and `deepseek-v4-flash-vision-exp` still work but are served by V4.1-Flash at its lower rate. The V4-Flash weights stay on Hugging Face. Still no consumer local path — see the Sep 10 fast take. **Sibling — August 31, 2026:** `deepseek-ai/DeepSeek-V4-Flash-Vision-Exp` put open MIT weights on HF for the vision variant (305B total, image-text-to-text, 209K downloads in its first week; hosted since Aug 21 as `deepseek-v4-flash-vision-exp` at V4-Flash rates, images billed at up to 384 tokens each). Same big-iron class as this entry — no local path, no pick change; see the Aug 31 fast take. A dated refresh landed on July 31, 2026 as `deepseek-ai/DeepSeek-V4-Flash-0731` — still MIT, still 48 safetensors shards, but 304B total rather than 284B, and it picked up 156K downloads in its first days. Treat it as the current checkpoint of this same entry, not a new model. Unsloth dynamic GGUFs at `unsloth/DeepSeek-V4-Flash` (early community quants mishandled the MoE router — use Unsloth's). Correction (September 2026): this note used to say an Ollama tag `deepseek-v4-flash` exists for local use. The Ollama library only lists `deepseek-v4-flash:cloud` and `:0731-cloud` — hosted tags, not downloadable weights (checked September 15, 2026). vLLM + multi-GPU is the cleanest production path. It also said we had no published rate card and so no cost-calculator row; that stopped being true in August — the calculator carries DeepSeek's Flash row, now at the V4.1-Flash rate.
Next step
Find-by-model — see what hardware runs this→