Google logo

Gemma 4 12B

by Google · Released · Apache 2.0

Mid-size dense model with native tool use, reasoning, and multimodal capability.

Parameters
12B dense
Context
256K tokens
Input
Text + image
Speculative decoding
Not yet
Intelligence
25%
License
Apache 2.0
Variants
Q4 QAT

Downloads

Sizes include the vision projector Magnitude installs. Every file is pinned to an immutable Hugging Face commit.

Variant Format Download Fidelity
Q4 QAT UD-Q4_K_XLQAT 6.9 GB 58/100

How it compares

Intelligence across every model in the catalog: each model's Artificial Analysis Intelligence Index score as a percentage of the top model on v4.3.2, as of Sep 29, 2026.

  1. Qwen3.8 27B 59%
  2. Gemma 4 31B 33%
  3. Qwen3.6 35B-A3B 32%
  4. Muse Glimmer 30B 30%
  5. Gemma 4 26B-A4B 29%
  6. Gemma 4 12B 25%
  7. Qwen3.5 4B 23%
  8. MiniCPM5 2B 22%
  9. Nemotron 3.5 Lightning 30B-A3B 22%
  10. Qwen3.5 9B 19%
  11. Gemma 4 E4B 15%
  12. Liquid LFM2.5 2.6B 15%
  13. MiniCPM5 1B 15%
  14. Gemma 4 E2B 14%
  15. Liquid LFM2.5 8B-A1B 13%

Works with your agent

One click in Magnitude connects Gemma 4 12B to any of these harnesses.

  • Pi
  • OpenCode
  • Hermes
  • OpenClaw
  • Codex
  • Claude Code
  • Oh My Pi
  • Cline

Related models

FAQ

How much memory does Gemma 4 12B need?

The smallest download is 6.9 GB. Plan for the weights plus a few gigabytes for context. Magnitude measures your free memory and only recommends variants that fit with room to work.

Can I run Gemma 4 12B on a Mac?

Yes. Magnitude runs GGUF models on Apple Silicon through Metal, and on NVIDIA, AMD, and CPU-only machines on Windows and Linux. The app estimates tokens per second for your exact chip before you download.

Which variant should I pick?

Magnitude picks for you: the highest-fidelity variant that fits your memory. If you choose manually, Q4 QAT keeps the most fidelity.

Catalog data from magnitudedev/magnitude, updated Sep 29, 2026.

Run Gemma 4 12B on your machine