the AI bench
VERIFIED SEPTEMBER 2026
All models

MODEL · DEEPSEEK · 1.65T TOTAL / 49B ACTIVE (MOE); 43 LAYERS, 256 ROUTED EXPERTS + 1 SHARED, 6 ACTIVE PER TOKEN, FP4 EXPERT WEIGHTS

DeepSeek V4-Pro

DeepSeek's frontier-class V4 flagship — 1.6T MoE that matches GPT-5.4 and Sonnet 4.6 on most benchmarks at meaningfully lower hosted price. The 1M-context default uses ~27% of V3.2's single-token FLOPs and ~10% of its KV cache thanks to architecture changes. MIT-licensed, but not realistically a local pick at this size.

License: MIT — and now the largest open-weight model under a classic OSI licence · Context: 1M tokens (384K max output) · Released: April 24, 2026 (preview); official release as DeepSeek-V4-Pro-0813 on August 13, 2026

The decision in five lines

The call
Hosted only
Best for
Hosted reference and benchmarks
Runs on
Hosted or workstation-class only · ~800 GB+ (not consumer-local; hosted only realistic)
Watch out
Use the DeepSeek API or a hosted route — Together, OpenRouter, or DeepSeek's own endpoints — and let the smaller V4-Flash sibling fill multi-GPU local roles.
Evidence
Estimated · last verified September 2026

1.65T total
PARAMETERS
MOE
TYPE
1M
CONTEXT
~800 GB+ (not consumer-local; hosted only realistic)
VRAM AT Q4

Where we recommend this

This model isn’t currently in an active planner slot. See the runner notes below if you’re running it anyway.

The call

DeepSeek's frontier-class V4 flagship — 1.6T MoE that matches GPT-5.4 and Sonnet 4.6 on most benchmarks at meaningfully lower hosted price. The 1M-context default uses ~27% of V3.2's single-token FLOPs and ~10% of its KV cache thanks to architecture changes. MIT-licensed, but not realistically a local pick at this size.

When not to use: Local hardware budgets under 8× H100 (~$200K). At Q4 the weights alone exceed 800 GB. Use the DeepSeek API or a hosted route — Together, OpenRouter, or DeepSeek's own endpoints — and let the smaller V4-Flash sibling fill multi-GPU local roles.

Runner notes

The preview is superseded: `deepseek-ai/DeepSeek-V4-Pro-0813` landed August 13, 2026 — MIT, 1.65T parameters, with a DSpark speculative-decoding module attached at layers 40–42 and FP4 expert weights. DeepSeek reports large gains over the preview on agentic work (Terminal Bench 2.1 72.1 → 87.9, DeepSWE 12.8 → 62.7, NL2Repo 38.5 → 61.5) putting it level with Kimi K3 and ahead of Opus 4.8 on several of those rows — all vendor-run, on DeepSeek's own harness, so treat them as claims rather than findings until independent numbers exist. DeepSeek now publishes a rate card, which we could not verify on earlier sweeps: `deepseek-v4-pro` at $0.435 per 1M input on a cache miss ($0.003625 cache hit) and $0.87 output; `deepseek-v4-flash` at $0.14 / $0.28. ⚠️ That changes on **August 16, 2026 at 16:00 UTC**, when DeepSeek moves to peak/off-peak billing — v4-pro goes to $0.66/$1.98 off-peak and $1.32/$3.96 at peak (01:00–04:00 and 06:00–10:00 UTC), which is an increase on today's flat rate, not the usual off-peak discount. We have deliberately kept both out of the cost calculator until that structure settles rather than publish a number with a three-day shelf life. Unsloth dynamic GGUFs exist but the practical home is the API.

License
MIT — and now the largest open-weight model under a classic OSI licence
Released
April 24, 2026 (preview); official release as DeepSeek-V4-Pro-0813 on August 13, 2026
Maker
DeepSeek

Next step

Find-by-model — see what hardware runs this