MODEL · OPENBMB · 8B (SIGLIP-400M + WHISPER-MEDIUM + CHATTTS-200M + QWEN2.5-7B)
MiniCPM-o 2.6
GPT-4o-class omnimodal 8B — vision, speech input, speech output, voice cloning in one model. End-to-end with full-duplex live streaming.
License: Apache 2.0 (commercial use requires registration questionnaire) · Context: 32K · Released: January 2025
The decision in five lines
- The call
- Skip for local — for voice
- Best for
- voice
- Runs on
- 23 hardware picks fit (cheapest: Intel Arc B580 12 GB · $249)
- Watch out
- Pure TTS or pure STT — the full multimodal stack is overkill.
- Evidence
- Estimated
- 8B (SigLip-400M + Whisper-medium + ChatTTS-200M + Qwen2.5-7B)
- PARAMETERS
- MULTIMODAL
- TYPE
- 32K
- CONTEXT
- ~7 GB (int4) / ~16 GB (FP16)
- VRAM AT Q4
Where we recommend this
Every tier slot in the planner where this model is a top or alternate pick. Pulled live from planner.js — when the planner refreshes, this table stays current.
The call
GPT-4o-class omnimodal 8B — vision, speech input, speech output, voice cloning in one model. End-to-end with full-duplex live streaming.
When not to use: Pure TTS or pure STT — the full multimodal stack is overkill. Use Kokoro or faster-whisper for single-purpose pipelines.
Runner notes
llama.cpp for CPU, vLLM for throughput, or Ollama (`openbmb/minicpm-o2.6:8b`). int4 variant fits ~7 GB VRAM. Commercial use requires filling OpenBMB's registration form — not fully frictionless. Sibling `MiniCPM-V-4.6` (May 2026, vision-only) is the lighter and newer alternative when you don't need voice I/O.
Hardware that fits
Every hardware pick whose memory fits this model at the quant we recommend. Sorted cheapest-first — the top row is your best-value fit. Click through for the full buyer’s guide.
- Intel Arc B580 12 GBPerfect · 1.5× 12 GB · $249–$299
- NVIDIA RTX 3060 12 GBPerfect · 1.5× 12 GB · $280–$400
- Minisforum UM890 ProPerfect · 2.9× 32 GB DDR5 (shared) · $463–$580 all-in
- RTX 5060 Ti 16 GBPerfect · 1.9× 16 GB · $560–$610
- AMD Radeon RX 9070 XTPerfect · 1.9× 16 GB · $649–$779
- AMD Radeon RX 7900 XTXPerfect · 2.9× 24 GB · $760 used / ~$1,500 new
- Mac Mini M4 16 GBGood · 1.3× 16 GB unified · $799 (new floor) / $499–$599 (eBay/residuals)
- NVIDIA RTX 3090 (used, single)Perfect · 2.9× 24 GB · $950–$1,200
- NVIDIA RTX 5070 TiPerfect · 1.9× 16 GB · $980–$1,300
- NVIDIA RTX 5080Perfect · 1.9× 16 GB · $999–$1,400
- MacBook Air M5 24 GBPerfect · 1.9× 24 GB unified · $1,299–$1,699
- Mac Mini M4 Pro 24 GBPerfect · 1.9× 24 GB unified · $1,399
- Dual RTX 3090 (used)Perfect · 5.8× 48 GB · $1,800–$2,500 all-in
- Framework Desktop (Ryzen AI Max+ 395)Perfect · 10.4× 128 GB unified · $1,999–$2,851
- NVIDIA RTX 4090Perfect · 2.9× 24 GB · $2,200–$2,800
- M5 Pro MacBook Pro 48 GBPerfect · 3.9× 48 GB unified · $2,599–$3,099
- NVIDIA RTX 5090Perfect · 3.9× 32 GB · $2,910–$4,300
- Mac Studio M4 Max 64 GBPerfect · 5.2× 64 GB unified · $3,199
- NVIDIA RTX A6000 (48 GB, used)Perfect · 5.8× 48 GB ECC · $3,500–$4,500
- Mac Studio M3 Ultra 96 GBPerfect · 7.8× 96 GB unified · $3,999
- M5 Max MacBook Pro 64 GBPerfect · 5.2× 64 GB unified · $4,499
- NVIDIA DGX SparkPerfect · 10.4× 128 GB unified · $4,699
- Dual RTX 5090Perfect · 7.7× 64 GB (2×32) · $8,500–$10,500
Next step
Find-by-model — see what hardware runs this→