DeepSeek V4 Flash

MIT

DeepSeek Β· 158B (13B active) Β· Mixture of Experts

Efficient long-context V4 β€” 13B active, 1M context Check if your GPU or Mac can run DeepSeek V4 Flash locally β€” 88.3 GB min, 147.1 GB recommended.

2026-041024K context

Mixture of Experts

Total experts: 256
Active experts: 6
Active params: 13.0B

Quantization Options

QuantBitsVRAMQualityStatus
Q2_K251.1 GBlowβ€”
Q3_K_M371.3 GBmoderateβ€”
Q4_K_M481.4 GBgoodβ€”
Q5_K_M5101.7 GBgoodβ€”
Q6_K6121.9 GBexcellentβ€”
Q8_08162.4 GBexcellentβ€”
F1616324.2 GBlosslessβ€”

Can I run DeepSeek V4 Flash locally?

Can I run DeepSeek V4 Flash locally?
DeepSeek V4 Flash needs about 88.3 GB of memory at a minimum and 147.1 GB recommended. Open this page to grade it against your GPU or Mac, then run it with runai, Ollama or LM Studio.
How much VRAM does DeepSeek V4 Flash need?
At Q4_K_M, DeepSeek V4 Flash uses about 81.4 GB of VRAM. Higher quants need more memory; lower quants fit tighter cards with a quality tradeoff.