Browse open models

Every model in the Magnitude catalog, tuned and tested for Apple Silicon, NVIDIA, AMD, and CPU.

15 models

Qwen logo

Qwen3.8 27B

Qwen ·

Large dense multimodal model with frontier-class agentic coding capability for its size.

  • 27B dense
  • 256K context
  • Vision
  • DFlash2
  • Apache 2.0

59% Intelligence

  • Q420.5 GB
  • Q628.3 GB
  • Q834.4 GB
Google logo

Gemma 4 31B

Google ·

Large dense Gemma model with the strongest measured coding score in its family.

  • 31B dense
  • 256K context
  • Vision
  • Apache 2.0

33% Intelligence

  • Q4 QAT18.5 GB
Qwen logo

Qwen3.6 35B-A3B

Qwen ·

Efficient MoE coding model with a large knowledge footprint and low active compute.

  • 35B · 3B active
  • 256K context
  • Vision
  • DFlash
  • Apache 2.0

32% Intelligence

  • Q424.2 GB
  • Q528.5 GB
  • Q633.9 GB
  • Q840.4 GB
Meta logo

Muse Glimmer 30B

Meta ·

Dense multimodal model purpose-built for autonomous agentic tasks on consumer hardware.

  • 30B dense
  • 128K context
  • Vision
  • DFlash
  • Apache 2.0

30% Intelligence

  • Q419.6 GB
  • Q525.5 GB
  • Q629.9 GB
  • Q836.0 GB
Google logo

Gemma 4 26B-A4B

Google ·

Mid-size MoE model balancing a substantial weight footprint with low active compute.

  • 25B · 3.8B active
  • 256K context
  • Vision
  • Apache 2.0

29% Intelligence

  • Q4 QAT15.4 GB
Google logo

Gemma 4 12B

Google ·

Mid-size dense model with native tool use, reasoning, and multimodal capability.

  • 12B dense
  • 256K context
  • Vision
  • Apache 2.0

25% Intelligence

  • Q4 QAT6.9 GB
Qwen logo

Qwen3.5 4B

Qwen ·

Compact dense model for machines where responsiveness and footprint matter most.

  • 4B dense
  • 256K context
  • Vision
  • Apache 2.0

23% Intelligence

  • Q43.7 GB
  • Q54.0 GB
  • Q64.9 GB
  • Q86.7 GB
OpenBMB logo

MiniCPM5 2B

OpenBMB ·

Compact dense model with native long context, tool use, and reasoning for on-device work.

  • 2.5B dense
  • 128K context
  • DSpark
  • Apache 2.0

22% Intelligence

  • Q42.2 GB
  • Q83.3 GB
NVIDIA logo

Nemotron 3.5 Lightning 30B-A3B

NVIDIA ·

Efficient hybrid MoE model for local reasoning, coding, and agentic workflows.

  • 30B · 3B active
  • 1M context
  • DFlash
  • OpenMDW 1.1

22% Intelligence

  • NVFP4 QAT23.6 GB
  • Q426.6 GB
  • Q836.2 GB
Qwen logo

Qwen3.5 9B

Qwen ·

Small dense model with a substantial capability gain over the 4B tier.

  • 9B dense
  • 256K context
  • Vision
  • Apache 2.0

19% Intelligence

  • Q47.1 GB
  • Q57.8 GB
  • Q69.9 GB
  • Q814.2 GB
Google logo

Gemma 4 E4B

Google ·

Small dense model with per-layer embeddings that materially improves capability over the E2B tier while remaining practical on memory-constrained machines.

  • 8B dense
  • 128K context
  • Vision
  • Apache 2.0

15% Intelligence

  • Q4 QAT5.2 GB
Liquid AI logo

Liquid LFM2.5 2.6B

Liquid AI ·

Compact dense Liquid model tuned for fast on-device tool use and coding workflows.

  • 2.6B dense
  • 128K context
  • DSpark
  • LFM 1.0

15% Intelligence

  • Q42.0 GB
  • Q52.3 GB
  • Q62.6 GB
  • Q83.2 GB
OpenBMB logo

MiniCPM5 1B

OpenBMB ·

Compact dense model for on-device reasoning and tool use.

  • 1.1B dense
  • 128K context
  • Apache 2.0

15% Intelligence

  • Q40.7 GB
  • Q81.2 GB
Google logo

Gemma 4 E2B

Google ·

Very small dense model with per-layer embeddings optimized for on-device use.

  • 5.1B dense
  • 128K context
  • Vision
  • Apache 2.0

14% Intelligence

  • Q4 QAT3.6 GB
Liquid AI logo

Liquid LFM2.5 8B-A1B

Liquid AI ·

Low-active-parameter Liquid MoE optimized for fast local reasoning and tool use.

  • 8.3B · 1.5B active
  • 125K context
  • DSpark
  • LFM 1.0

13% Intelligence

  • Q45.5 GB
  • Q56.4 GB
  • Q67.3 GB
  • Q89.4 GB

How to read this catalog

What is the intelligence score?

How a model compares with the most capable model available. It is the model’s Artificial Analysis Intelligence Index score, a composite of reasoning, coding, math, and knowledge benchmarks, as a percentage of the top model on the index. Where Artificial Analysis has not fully measured a model, Magnitude uses an estimate.

What do Q4, Q5, Q6, and Q8 mean?

Quantization levels. Lower numbers mean smaller downloads and less memory at a small cost in fidelity. Variants marked QAT were trained with quantization in mind and keep more fidelity at the same size. Magnitude recommends the highest-fidelity variant that fits your memory with room for context.

How are download sizes measured?

Weights plus every companion file Magnitude installs: the vision projector for image-capable models and the draft model for speculative decoding. They are the exact files pinned by the catalog’s locked Hugging Face commits.

Will these run on my machine?

Magnitude profiles your hardware and estimates tokens per second for every model before you download. The sizes here are what you download; the app gives you the real answer for your chip and memory, including room for context.

Catalog data from magnitudedev/magnitude, updated Sep 29, 2026.

See what your machine can run