MODEL · DEEPSEEK · 1.65T TOTAL / 49B ACTIVE (MOE); 43 LAYERS, 256 ROUTED EXPERTS + 1 SHARED, 6 ACTIVE PER TOKEN, FP4 EXPERT WEIGHTS
DeepSeek V4-Pro
DeepSeek's frontier-class V4 flagship — 1.6T MoE that matches GPT-5.4 and Sonnet 4.6 on most benchmarks at meaningfully lower hosted price. The 1M-context default uses ~27% of V3.2's single-token FLOPs and ~10% of its KV cache thanks to architecture changes. MIT-licensed, but not realistically a local pick at this size.
License: MIT — and now the largest open-weight model under a classic OSI licence · Context: 1M tokens (384K max output) · Released: April 24, 2026 (preview); official release as DeepSeek-V4-Pro-0813 on August 13, 2026
The decision in five lines
- The call
- Hosted only
- Best for
- Hosted reference and benchmarks
- Runs on
- Hosted or workstation-class only · ~800 GB+ (not consumer-local; hosted only realistic)
- Watch out
- Use the DeepSeek API or a hosted route — Together, OpenRouter, or DeepSeek's own endpoints — and let the smaller V4-Flash sibling fill multi-GPU local roles.
- Evidence
- Estimated
- 1.65T total
- PARAMETERS
- MOE
- TYPE
- 1M
- CONTEXT
- ~800 GB+ (not consumer-local; hosted only realistic)
- VRAM AT Q4
Where we recommend this
This model isn’t currently in an active planner slot. See the runner notes below if you’re running it anyway.
The call
DeepSeek's frontier-class V4 flagship — 1.6T MoE that matches GPT-5.4 and Sonnet 4.6 on most benchmarks at meaningfully lower hosted price. The 1M-context default uses ~27% of V3.2's single-token FLOPs and ~10% of its KV cache thanks to architecture changes. MIT-licensed, but not realistically a local pick at this size.
When not to use: Local hardware budgets under 8× H100 (~$200K). At Q4 the weights alone exceed 800 GB. Use the DeepSeek API or a hosted route — Together, OpenRouter, or DeepSeek's own endpoints — and let the smaller V4-Flash sibling fill multi-GPU local roles.
Runner notes
The preview is superseded: `deepseek-ai/DeepSeek-V4-Pro-0813` landed August 13, 2026 — MIT, 1.65T parameters, with a DSpark speculative-decoding module attached at layers 40–42 and FP4 expert weights. DeepSeek reports large gains over the preview on agentic work (Terminal Bench 2.1 72.1 → 87.9, DeepSWE 12.8 → 62.7, NL2Repo 38.5 → 61.5) putting it level with Kimi K3 and ahead of Opus 4.8 on several of those rows — all vendor-run, on DeepSeek's own harness, so treat them as claims rather than findings until independent numbers exist. DeepSeek now publishes a rate card, which we could not verify on earlier sweeps: `deepseek-v4-pro` at $0.435 per 1M input on a cache miss ($0.003625 cache hit) and $0.87 output; `deepseek-v4-flash` at $0.14 / $0.28. ⚠️ That changes on **August 16, 2026 at 16:00 UTC**, when DeepSeek moves to peak/off-peak billing — v4-pro goes to $0.66/$1.98 off-peak and $1.32/$3.96 at peak (01:00–04:00 and 06:00–10:00 UTC), which is an increase on today's flat rate, not the usual off-peak discount. We have deliberately kept both out of the cost calculator until that structure settles rather than publish a number with a three-day shelf life. Unsloth dynamic GGUFs exist but the practical home is the API.
Next step
Find-by-model — see what hardware runs this→