A newer version is available: GLM-5.3View β†’

GLM-5.1

MIT

Zhipu AI Β· 754B (40B active) Β· Mixture of Experts

Improved agentic coding β€” SOTA SWE-bench Pro, long-horizon tasks Check if your GPU or Mac can run GLM-5.1 locally β€” 421.3 GB min, 702.2 GB recommended.

2026-04128K context

Mixture of Experts

Total experts: 256
Active experts: 8
Active params: 40.0B

Quantization Options

QuantBitsVRAMQualityStatus
Q2_K2241.9 GBlowβ€”
Q3_K_M3338.4 GBmoderateβ€”
Q4_K_M4386.7 GBgoodβ€”
Q5_K_M5483.3 GBgoodβ€”
Q6_K6579.8 GBexcellentβ€”
Q8_08772.9 GBexcellentβ€”
F16161545.4 GBlosslessβ€”

Can I run GLM-5.1 locally?

Can I run GLM-5.1 locally?
GLM-5.1 needs about 421.3 GB of memory at a minimum and 702.2 GB recommended. Open this page to grade it against your GPU or Mac, then run it with runai, Ollama or LM Studio.
How much VRAM does GLM-5.1 need?
At Q4_K_M, GLM-5.1 uses about 386.7 GB of VRAM. Higher quants need more memory; lower quants fit tighter cards with a quality tradeoff.