42 models at Q4 · 8 GB budget

Local AI models for 8GB VRAM

8 GB is the most common consumer GPU budget. At Q4 you can run 3B–9B chat models comfortably, plus the smallest image checkpoints. Mixture-of-experts still load every expert, so big MoE files will not fit.

Typical fit: 3B–9B chat models at Q4, plus tiny image models. · Example:RTX 4060

These sizes assume Q4_K_M plus a little runtime overhead.Grade the full catalog on your machineorpick a specific GPU.

Common questions

What LLM can I run on 8GB VRAM?
7B–9B dense models at Q4 are the sweet spot. Smaller 1B–4B models will feel faster. Avoid 14B+ dense checkpoints and large MoE files unless you offload to system RAM (much slower).
Is 8GB enough for local image generation?
Yes for compact models such as FLUX Klein 4B or similar. Full FLUX.2 Dev and large video models will not fit.