OpenBMB logo

MiniCPM5 2B

by OpenBMB · Released · Apache 2.0

Compact dense model with native long context, tool use, and reasoning for on-device work.

Parameters
2.5B dense
Context
128K tokens
Input
Text
Speculative decoding
DSpark
Intelligence
22%
License
Apache 2.0
Variants
Q4 · Q8

Downloads

Sizes include the speculative draft Magnitude installs. Every file is pinned to an immutable Hugging Face commit.

Variant Format Download Fidelity
Q4 Q4_K_M 2.2 GB 40/100
Q8 Q8_0 3.3 GB 80/100

How it compares

Intelligence across every model in the catalog: each model's Artificial Analysis Intelligence Index score as a percentage of the top model on v4.3.2, as of Sep 29, 2026.

  1. Qwen3.8 27B 59%
  2. Gemma 4 31B 33%
  3. Qwen3.6 35B-A3B 32%
  4. Muse Glimmer 30B 30%
  5. Gemma 4 26B-A4B 29%
  6. Gemma 4 12B 25%
  7. Qwen3.5 4B 23%
  8. MiniCPM5 2B 22%
  9. Nemotron 3.5 Lightning 30B-A3B 22%
  10. Qwen3.5 9B 19%
  11. Gemma 4 E4B 15%
  12. Liquid LFM2.5 2.6B 15%
  13. MiniCPM5 1B 15%
  14. Gemma 4 E2B 14%
  15. Liquid LFM2.5 8B-A1B 13%

Works with your agent

One click in Magnitude connects MiniCPM5 2B to any of these harnesses.

  • Pi
  • OpenCode
  • Hermes
  • OpenClaw
  • Codex
  • Claude Code
  • Oh My Pi
  • Cline

Related models

FAQ

How much memory does MiniCPM5 2B need?

The smallest download is 2.2 GB and the largest is 3.3 GB. Plan for the weights plus a few gigabytes for context. Magnitude measures your free memory and only recommends variants that fit with room to work.

Can I run MiniCPM5 2B on a Mac?

Yes. Magnitude runs GGUF models on Apple Silicon through Metal, and on NVIDIA, AMD, and CPU-only machines on Windows and Linux. The app estimates tokens per second for your exact chip before you download.

Which variant should I pick?

Magnitude picks for you: the highest-fidelity variant that fits your memory. If you choose manually, Q8 keeps the most fidelity, and Q4 is the smallest at 2.2 GB.

Catalog data from magnitudedev/magnitude, updated Sep 29, 2026.

Run MiniCPM5 2B on your machine