MODEL · OPENBMB · 1.08B TOTAL / 679M NON-EMBEDDING (LLAMAFORCAUSALLM)
MiniCPM5-1B (Apache 2.0, OPD-trained)
OpenBMB's claimed 1B-class open-source SOTA — but the training-method story matters more than the size. The post-training pipeline runs SFT → RL → On-Policy Distillation (OPD): RL teachers are trained per domain (math, code, closed-book QA, writing) and then distilled back into one release model. RL + OPD lifts the SFT-only checkpoint by +16pt average on math / code / instruction-following and drops max-token-truncated responses by 29 percentage points. Hybrid `<think>` reasoning toggle (switch via `enable_thinking`) and native XML-style tool calling. English + Chinese.
License: Apache 2.0 · Context: 131,072 (128K) · Released: May 22, 2026
The decision in five lines
- The call
- Consider — runnable locally, family reference
- Best for
- Local evaluation and family reference
- Runs on
- 23 hardware picks fit (cheapest: Intel Arc B580 12 GB · $310)
- Watch out
- General-purpose chat at any size larger than 4B — Qwen 3.5 4B, Phi-4 Mini, and Gemma 3 4B all win on absolute capability when you can fit them.
- Evidence
- Estimated
- 1.08B total
- PARAMETERS
- DENSE LLM
- TYPE
- 131072
- CONTEXT
- ~0.7–1 GB (Q4) / ~2.2 GB (FP16)
- VRAM AT Q4
Where we recommend this
This model isn’t currently in an active planner slot. See the runner notes below if you’re running it anyway.
The call
OpenBMB's claimed 1B-class open-source SOTA — but the training-method story matters more than the size. The post-training pipeline runs SFT → RL → On-Policy Distillation (OPD): RL teachers are trained per domain (math, code, closed-book QA, writing) and then distilled back into one release model. RL + OPD lifts the SFT-only checkpoint by +16pt average on math / code / instruction-following and drops max-token-truncated responses by 29 percentage points. Hybrid `<think>` reasoning toggle (switch via `enable_thinking`) and native XML-style tool calling. English + Chinese.
When not to use: General-purpose chat at any size larger than 4B — Qwen 3.5 4B, Phi-4 Mini, and Gemma 3 4B all win on absolute capability when you can fit them. MiniCPM5-1B wins on "smallest credible model that thinks and calls tools."
Runner notes
Standard `LlamaForCausalLM` weights — llama.cpp / Ollama / vLLM paths all work day-one. SGLang is the recommended backend for tool calling. Sibling SFT-only checkpoint at `OpenBMB/MiniCPM5-1B-SFT` for ablation. Recommended sampling: think mode `temperature=0.9, top_p=0.95, enable_thinking=True`; no-think `temperature=0.7, top_p=0.95`. Sits parallel to BitCPM4-CANN family — both OpenBMB sub-2B research milestones, but different angles (BitCPM4 is native ternary architecture, MiniCPM5 is OPD post-training). **2B sibling — September 6, 2026:** `openbmb/MiniCPM5-2B` (Apache 2.0, 2.52B total / 1.98B non-embedding, 42 layers, GQA 16/2, 131K context) scales the same recipe up, with official GGUF, MLX 4-bit, GPTQ and a DSpark draft model; 206K downloads in nine days. OpenBMB claims 2B-class SOTA on its own comparison set — vendor numbers. No Ollama library tag; use the official GGUF, and raise llama.cpp's default `min_p=0.05`, which OpenBMB says can cause repetition. See the Sep 6 fast take.
Hardware that fits
Every hardware pick whose memory fits this model at the quant we recommend. Sorted cheapest-first — the top row is your best-value fit. Click through for the full buyer’s guide.
- Intel Arc B580 12 GBPerfect · 6.6× 12 GB · $310–$429 new / $190–$260 refurb
- Minisforum UM890 ProPerfect · 13.2× 32 GB DDR5 (shared) · $463 barebone (no RAM / SSD); $991 with 32 GB / 1 TB — Minisforum US store, Sep 15 2026
- NVIDIA RTX 3060 12 GBPerfect · 6.6× 12 GB · $480–$660
- RTX 5060 Ti 16 GBPerfect · 8.8× 16 GB · $700 (refurb / open box) – $970 (new)
- AMD Radeon RX 9070 XTPerfect · 8.8× 16 GB · $750–$870
- Mac Mini M4 16 GBPerfect · 5.9× 16 GB unified · $899 (M6 successor, 16 GB / 256 GB, pre-order) / M4 residuals $499–$799
- NVIDIA RTX 3090 (used, single)Perfect · 13.2× 24 GB · $950–$1,200
- NVIDIA RTX 5070 TiPerfect · 8.8× 16 GB · $1,150–$1,300 new; open-box / refurb from ~$1,020
- AMD Radeon RX 7900 XTXPerfect · 13.2× 24 GB · $1,300–$1,500 new (no refurb in stock Sep 15 2026)
- NVIDIA RTX 5080Perfect · 8.8× 16 GB · $1,430–$1,700
- MacBook Air M5 24 GBPerfect · 8.8× 24 GB unified · $1,499–$1,899
- Mac Mini M4 Pro 24 GBPerfect · 8.8× 24 GB unified · $1,699 (M5 Pro successor, 24 GB / 512 GB, pre-order; M4 Pro discontinued Aug 25 2026)
- Dual RTX 3090 (used)Perfect · 26.3× 48 GB · $1,800–$2,500 all-in
- NVIDIA RTX 4090Perfect · 13.2× 24 GB · $2,200–$2,800
- M5 Pro MacBook Pro 48 GBPerfect · 17.6× 48 GB unified · $2,999–$3,599
- Framework Desktop (Ryzen AI Max+ 395)Perfect · 47.0× 128 GB unified · $3,449 (128 GB config)
- NVIDIA RTX A6000 (48 GB, used)Perfect · 26.3× 48 GB ECC · $3,500–$4,500
- Mac Studio M4 Max 64 GBPerfect · 23.5× 64 GB unified · $3,799 (M5 Max successor, 64 GB / 1 TB, pre-order; M4 Max discontinued Aug 25 2026)
- NVIDIA DGX SparkPerfect · 47.0× 128 GB unified · $4,699
- M5 Max MacBook Pro 64 GBPerfect · 23.5× 64 GB unified · ~$5,199 (est.; June 25 2026 increase)
- Mac Studio M3 Ultra 96 GBPerfect · 35.2× 96 GB unified · $5,499 (M5 Ultra successor, 96 GB / 1 TB, pre-order; M3 Ultra discontinued Aug 25 2026)
- NVIDIA RTX 5090Perfect · 17.5× 32 GB · $6,450–$7,000 (new, in stock)
- Dual RTX 5090Perfect · 35.1× 64 GB (2×32) · $13,300–$14,000 all-in (two cards alone are ~$12,900 at the Sep 15 2026 Newegg floor)
Next step
Find-by-model — see what hardware runs this→