MiniCPM5 2B
by OpenBMB · Released · Apache 2.0
Compact dense model with native long context, tool use, and reasoning for on-device work.
- Parameters
- 2.5B dense
- Context
- 128K tokens
- Input
- Text
- Speculative decoding
- DSpark
- Intelligence
- 22%
- License
- Apache 2.0
- Variants
- Q4 · Q8
- Source
- Hugging Face ↗
Downloads
Sizes include the speculative draft Magnitude installs. Every file is pinned to an immutable Hugging Face commit.
| Variant | Format | Download | Fidelity |
|---|---|---|---|
| Q4 | Q4_K_M | 2.2 GB | 40/100 |
| Q8 | Q8_0 | 3.3 GB | 80/100 |
How it compares
Intelligence across every model in the catalog: each model's Artificial Analysis Intelligence Index score as a percentage of the top model on v4.3.2, as of Sep 29, 2026.
- Qwen3.8 27B 59%
- Gemma 4 31B 33%
- Qwen3.6 35B-A3B 32%
- Muse Glimmer 30B 30%
- Gemma 4 26B-A4B 29%
- Gemma 4 12B 25%
- Qwen3.5 4B 23%
- MiniCPM5 2B 22%
- Nemotron 3.5 Lightning 30B-A3B 22%
- Qwen3.5 9B 19%
- Gemma 4 E4B 15%
- Liquid LFM2.5 2.6B 15%
- MiniCPM5 1B 15%
- Gemma 4 E2B 14%
- Liquid LFM2.5 8B-A1B 13%
Works with your agent
One click in Magnitude connects MiniCPM5 2B to any of these harnesses.
-
Pi
-
OpenCode
-
Hermes
-
OpenClaw
-
Codex
-
Claude Code
-
Oh My Pi
-
Cline
Related models
FAQ
How much memory does MiniCPM5 2B need?
The smallest download is 2.2 GB and the largest is 3.3 GB. Plan for the weights plus a few gigabytes for context. Magnitude measures your free memory and only recommends variants that fit with room to work.
Can I run MiniCPM5 2B on a Mac?
Yes. Magnitude runs GGUF models on Apple Silicon through Metal, and on NVIDIA, AMD, and CPU-only machines on Windows and Linux. The app estimates tokens per second for your exact chip before you download.
Which variant should I pick?
Magnitude picks for you: the highest-fidelity variant that fits your memory. If you choose manually, Q8 keeps the most fidelity, and Q4 is the smallest at 2.2 GB.
Catalog data from magnitudedev/magnitude, updated Sep 29, 2026.