FAST TAKE · 2026-08-31 · DEEPSEEK-V4-FLASH-VISION-EXP
DeepSeek-V4-Flash-Vision-Exp: the V4-Flash MoE learns to see, and the weights are MIT
DeepSeek put the weights of DeepSeek-V4-Flash-Vision-Exp on HuggingFace on August 31, 2026 under the MIT licence: 305B parameters in safetensors, image-text-to-text, and 209,191 downloads by September 7. The hosted version had been live on the DeepSeek API since August 21 as `deepseek-v4-flash-vision-exp`. DeepSeek's own claim is that it "matches DeepSeek-V4-Flash on text capabilities" while "bringing multimodal agent performance close to Opus-4.8" — a vendor benchmark statement we pass along as exactly that.
Verdict: DeepSeek's first open-weight vision model — MIT, 305B total, 209K downloads in a week, and still big-iron
The take
The facts from the card and the API notice. Weights: `deepseek-ai/DeepSeek-V4-Flash-Vision-Exp`, created 2026-08-31, MIT, 304,646,824,126 parameters in mixed BF16 / F8 / I8 tensors, served with SGLang using the same DSpark speculative decoding as the text model. API: images are tokenized for billing at up to 384 tokens each, charged at V4-Flash rates; Chat Completions, Messages and Responses endpoints all accept mixed text + image input, and a free Files API lets you upload an image once and reference it by id. DeepSeek Harness 0.1.1 shipped the same day with support built in.
What it is for local buyers: nothing yet, and honestly so. At 305B total this is the same class as the V4-Flash text model it extends — roughly 150 GB even at Q4 by our arithmetic, past every box on our hardware list, and there are no official GGUF or MLX builds. The MoE's small active set keeps per-token compute low, but the bill is memory, and no consumer machine has it. It stays in the hosted-comparator bucket alongside V4-Flash, Kimi K3 and GLM-5.3-Flash.
Why it still earns a dated verdict: it is the first open-weight, permissively licensed vision model from DeepSeek, and 209K downloads in a week says the ecosystem treats it as real rather than experimental despite the "Exp" suffix. If DeepSeek follows the pattern it set with the text line — a dated refresh, then quantized community builds — the interesting question becomes whether a distilled or smaller sibling lands in the 24–48 GB band, where Qwen3.8-27B currently owns local vision. We will look again when one does.
Our call: no model page, no planner-pick change. The `/models/deepseek-v4-flash/` entry now notes this sibling. Hosted users get the multimodal route at the existing V4-Flash rate card, which remains on our permanent hold for clock-dependent pricing.
Where this fits
Models: DeepSeek V4-Flash · Qwen3.8-27B · Mage-VL (4B codec-native video VLM)
Hardware: NVIDIA RTX 5090 · Mac Studio M3 Ultra 96 GB
Sources
Next step
Try this in the planner→