Gemma 4 E2B
by Google · Released · Apache 2.0
Very small dense model with per-layer embeddings optimized for on-device use.
- Parameters
- 5.1B dense
- Context
- 128K tokens
- Input
- Text + image
- Speculative decoding
- Not yet
- Intelligence
- 14%
- License
- Apache 2.0
- Variants
- Q4 QAT
- Source
- Hugging Face ↗
Downloads
Sizes include the vision projector Magnitude installs. Every file is pinned to an immutable Hugging Face commit.
| Variant | Format | Download | Fidelity |
|---|---|---|---|
| Q4 QAT | UD-Q4_K_XLQAT | 3.6 GB | 58/100 |
How it compares
Intelligence across every model in the catalog: each model's Artificial Analysis Intelligence Index score as a percentage of the top model on v4.3.2, as of Sep 29, 2026.
- Qwen3.8 27B 59%
- Gemma 4 31B 33%
- Qwen3.6 35B-A3B 32%
- Muse Glimmer 30B 30%
- Gemma 4 26B-A4B 29%
- Gemma 4 12B 25%
- Qwen3.5 4B 23%
- MiniCPM5 2B 22%
- Nemotron 3.5 Lightning 30B-A3B 22%
- Qwen3.5 9B 19%
- Gemma 4 E4B 15%
- Liquid LFM2.5 2.6B 15%
- MiniCPM5 1B 15%
- Gemma 4 E2B 14%
- Liquid LFM2.5 8B-A1B 13%
Works with your agent
One click in Magnitude connects Gemma 4 E2B to any of these harnesses.
-
Pi
-
OpenCode
-
Hermes
-
OpenClaw
-
Codex
-
Claude Code
-
Oh My Pi
-
Cline
Related models
FAQ
How much memory does Gemma 4 E2B need?
The smallest download is 3.6 GB. Plan for the weights plus a few gigabytes for context. Magnitude measures your free memory and only recommends variants that fit with room to work.
Can I run Gemma 4 E2B on a Mac?
Yes. Magnitude runs GGUF models on Apple Silicon through Metal, and on NVIDIA, AMD, and CPU-only machines on Windows and Linux. The app estimates tokens per second for your exact chip before you download.
Which variant should I pick?
Magnitude picks for you: the highest-fidelity variant that fits your memory. If you choose manually, Q4 QAT keeps the most fidelity.
Catalog data from magnitudedev/magnitude, updated Sep 29, 2026.