Liquid AI Β· 24B (2.3B active) Β· Mixture of Experts

Hybrid MoE with convolution+attention layers β€” 2.3B active Check if your GPU or Mac can run LFM2 24B locally β€” 13.4 GB min, 22.4 GB recommended.

2025-1132K context

Mixture of Experts

Total experts: 64
Active experts: 4
Active params: 2.3B

Quantization Options

QuantBitsVRAMQualityStatus
Q2_K28.2 GBlowβ€”
Q3_K_M311.3 GBmoderateβ€”
Q4_K_M412.8 GBgoodβ€”
Q5_K_M515.9 GBgoodβ€”
Q6_K618.9 GBexcellentβ€”
Q8_0825.1 GBexcellentβ€”
F161649.7 GBlosslessβ€”

About this model

image.png

LFM2 is a family of hybrid models designed for on-device deployment. LFM2-24B-A2B is the largest model in the family, scaling the architecture to 24 billion parameters while keeping inference efficient.

  • Best-in-class efficiency: A 24B MoE model with only 2B active parameters per token, fitting in 32 GB of RAM for deployment on consumer laptops and desktops.
  • Fast edge inference: 112 tok/s decode on AMD CPU, 293 tok/s on H100. Fits in 32B GB of RAM.
  • Predictable scaling: Quality improves log-linearly from 350M to 24B total parameters, confirming the LFM2 hybrid architecture scales reliably across nearly two orders of magnitude.

image.png

Can I run LFM2 24B locally?

Can I run LFM2 24B locally?
LFM2 24B needs about 13.4 GB of memory at a minimum and 22.4 GB recommended. Open this page to grade it against your GPU or Mac, then run it with runai, Ollama or LM Studio.
How much VRAM does LFM2 24B need?
At Q4_K_M, LFM2 24B uses about 12.8 GB of VRAM. Higher quants need more memory; lower quants fit tighter cards with a quality tradeoff.