FLUX.2 Klein 4B
Apache 2.0Black Forest Labs Β· 4B Β· Apache 2.0
Fastest open FLUX.2 β sub-second text-to-image and multi-reference editing on consumer GPUs
We detect your GPU or Mac and rank the open models that actually fit.
Top open models for your device
Chat, code and reasoning
Generate images locally
Generate video locally
On-device, 4B and under
This device is very constrained
Nothing runs comfortably yet. Switch to Browse all to see the lightest options.
Black Forest Labs Β· 4B Β· Apache 2.0
Fastest open FLUX.2 β sub-second text-to-image and multi-reference editing on consumer GPUs
Alibaba Β· 5B Β· Apache 2.0
Unified text/image-to-video β the local sweet spot under Apache 2.0
Alibaba Β· 6B Β· Apache 2.0
8-step distilled image model β photorealism and bilingual text on 16GB cards
Tencent Β· 8.3B Β· Tencent Hunyuan Community
Compact cinematic video model β strong faces and motion on a single 4090
Alibaba Β· 8.8B Β· Apache 2.0
The community-favourite local VLM β superb OCR, receipts & captioning
Alibaba Β· 9B Β· Apache 2.0
Multimodal Qwen 3.5 mid-size
Lightricks Β· 19B Β· LTX-2 Community
Open 4K video with native stereo audio β text, image and video-to-video
Alibaba Β· 20B Β· Apache 2.0
Open text-to-image with strong English and Chinese typography
OpenAI Β· 21B Β· Apache 2.0
OpenAI's open-weight MoE with configurable reasoning
Google Β· 27B Β· Gemma
Gemma 4 MoE instruct model (official)
Alibaba Β· 27B Β· Apache 2.0
Flagship dense Qwen 3.8 β native multimodal all-rounder with video understanding
Alibaba Β· 27B Β· Apache 2.0
Flagship open Wan 2.2 β 14B-active MoE for photoreal text-to-video
Meta Β· 30B Β· Apache 2.0
Open agentic 30B distilled from Muse Spark β tool use, vision and local recovery on a single GPU
Alibaba Β· 31B Β· Apache 2.0
Efficient vision MoE β 3B active, strong temporal & document understanding
Black Forest Labs Β· 32B Β· FLUX Non-Commercial
Flagship open-weight FLUX.2 β text-to-image and multi-reference editing up to 4MP
MiniMax Β· 33B Β· MiniMax Community
Open video generation β text/image to 2K video with native stereo audio
InternScience Β· 35B Β· Apache 2.0
Efficient multimodal agentic MoE for long-horizon search, engineering and scientific research
DeepReinforce Β· 35B Β· MIT
Agentic coding MoE with a 3B active working set and self-improving training
Alibaba Β· 36B Β· Apache 2.0
Big-model quality at 3B-active speed β the mid-hardware sweet spot
Tencent Β· 80B Β· Tencent Hunyuan Community
Reasoning image model β prompt rewrite, chain-of-thought and image-to-image editing
Meta Β· 109B Β· Llama 4 Community
MoE with 16 experts, 17B active params
OpenAI Β· 117B Β· Apache 2.0
OpenAI's flagship open-weight MoE β 52.6% SWE-bench
Mistral AI Β· 119B Β· Apache 2.0
Sparse Mistral Small 4 β 6.5B active, strong local all-rounder
Alibaba Β· 235B Β· Apache 2.0
Flagship vision-language MoE β frontier multimodal reasoning and agentic GUI control
Tencent Β· 295B Β· Apache 2.0
Production-focused agentic MoE with strong coding, tool use and long-context reasoning
MiniMax Β· 428B Β· MiniMax Community
Native multimodal MoE β understands text, image and long video with 1M context
Z.ai Β· 753B Β· MIT
Same 753B / 40B-active base as GLM-5.2 β post-training lifts coding and long-horizon agents, 1M context
Meituan Β· 1.6T Β· MIT
Frontier-scale agentic and coding MoE with sparse attention and native 1M context
DeepSeek Β· 1.6T Β· MIT
Flagship V4 MoE β 49B active, 1M context
Alibaba Β· 2.4T Β· Qwen
Frontier Qwen 3.8 MoE β 95B active, 1M context
Moonshot AI Β· 2.8T Β· Kimi
Frontier 2.8T multimodal MoE β 104B active, native video understanding, 1M context
Alibaba Β· 0.6B Β· Apache 2.0
Ultra-light Qwen 3 model for constrained devices
Alibaba Β· 0.8B Β· Apache 2.0
Ultra-tiny model for embedded and edge
Meta Β· 1B Β· Llama 3.2 Community
Meta's smallest Llama for edge devices
Google Β· 1B Β· Gemma
Google's tiny Gemma for on-device
Alibaba Β· 1.3B Β· Apache 2.0
Tiny open text-to-video β 480p clips on 8GB consumer GPUs
Alibaba Β· 1.5B Β· Apache 2.0
Ultra-lightweight coding model
DeepSeek Β· 1.5B Β· MIT
Tiny reasoning model distilled from R1
Alibaba Β· 1.7B Β· Apache 2.0
Compact multilingual Qwen 3
Alibaba Β· 2B Β· Apache 2.0
Small multimodal Qwen 3.5
Meta Β· 3B Β· Llama 3.2 Community
Lightweight Llama for mobile and edge
HuggingFace Β· 3B Β· Apache 2.0
Lightweight multilingual reasoning
IBM Β· 3B Β· Apache 2.0
Compact enterprise model for edge and constrained environments
Mistral AI Β· 3B Β· Apache 2.0
Current-gen tiny Ministral β edge chat with 256K context
Microsoft Β· 3.8B Β· MIT
Lightweight reasoning model
Google Β· 4B Β· Gemma
Multimodal Gemma with 128K context
Alibaba Β· 4B Β· Apache 2.0
Small multimodal Qwen 3.5
Alibaba Β· 4.4B Β· Apache 2.0
Compact dedicated vision-language model β OCR & image chat on edge
Google Β· 5B Β· Gemma
Gemma 4 efficient instruct model (official)
Alibaba Β· 7B Β· Apache 2.0
Dedicated coding model
DeepSeek Β· 7B Β· MIT
R1 reasoning distilled into Qwen 7B
Google Β· 8B Β· Gemma
Gemma 4 balanced instruct model (official)
Meta Β· 8B Β· Llama 3.1 Community
Meta's versatile 8B β great quality/speed ratio
Alibaba Β· 8B Β· Apache 2.0
Qwen 3 with thinking mode support
IBM Β· 8B Β· Apache 2.0
Balanced general-purpose enterprise model
Mistral AI Β· 8B Β· MRL
Mistral's efficient 8B model
Zhipu AI Β· 9B Β· GLM-4
Multilingual model supporting 26 languages with 128K context
NVIDIA Β· 9B Β· NVIDIA Open
Hybrid Mamba2 architecture for reasoning
DeepReinforce Β· 9B Β· MIT
Self-improving agentic coding model optimized for terminal and software engineering tasks
Black Forest Labs Β· 9B Β· FLUX Non-Commercial
Higher-quality distilled FLUX.2 β sub-second generation and multi-reference editing
Google Β· 12B Β· Gemma
Multimodal Gemma with 128K context
Mistral AI Β· 12B Β· Apache 2.0
Multilingual 12B with 128K context
Google Β· 12B Β· Apache 2.0
Gemma 4 mid-size instruct β multimodal any-to-any
Microsoft Β· 14B Β· MIT
Microsoft's reasoning-focused model
Alibaba Β· 14B Β· Apache 2.0
Strong all-rounder with thinking mode
DeepSeek Β· 14B Β· MIT
R1 reasoning distilled into Qwen 14B
Mistral AI Β· 14B Β· Apache 2.0
Current-gen Ministral mid-size β local assistant with 256K context
Liquid AI Β· 24B Β· Liquid AI
Hybrid MoE with convolution+attention layers β 2.3B active
Mistral AI Β· 24B Β· Apache 2.0
Coding-focused model with 256K context β 68% SWE-bench
Mistral AI Β· 24B Β· Apache 2.0
Multimodal Mistral with vision support
Google Β· 26B Β· Apache 2.0
Discrete diffusion MoE β 1100+ tok/s on H100, multimodal (text/image/video)
Alibaba Β· 27.8B Β· Apache 2.0
Flagship native multimodal Qwen 3.5
Alibaba Β· 27.8B Β· Apache 2.0
Flagship dense Qwen 3.6 β native multimodal all-rounder
Alibaba Β· 30B Β· Apache 2.0
MoE with only 3.3B active β extremely efficient
NVIDIA Β· 30B Β· NVIDIA Open
MoE with 1M context and 3B active
IBM Β· 30B Β· Apache 2.0
High-capacity enterprise model for complex reasoning and tool use
Cohere Β· 30B Β· Apache 2.0
Open agentic coding MoE with 3B active β built for software engineering and terminal tasks
Alibaba Β· 30B Β· Apache 2.0
Efficient agentic coding MoE β 3B active, 256K context
Alibaba Β· 32B Β· Apache 2.0
Qwen 3 flagship dense model
DeepSeek Β· 32B Β· MIT
R1 reasoning distilled into Qwen 32B β sweet spot
Allen AI Β· 32B Β· Apache 2.0
Fully open research model by Allen AI
Google Β· 33B Β· Gemma
Gemma 4 flagship instruct model (official)
Google Β· 33B Β· Gemma
Gemma 4 flagship base model (official)
Cohere Β· 35B Β· CC BY-NC 4.0
Optimized for retrieval-augmented generation
Alibaba Β· 35B Β· Apache 2.0
Efficient multimodal MoE with 3B active
Mistral AI Β· 47B Β· Apache 2.0
MoE with 12.9B active params
Meta Β· 70B Β· Llama 3.3 Community
Best open model at 70B class
Alibaba Β· 80B Β· Apache 2.0
High-sparsity MoE β extreme low activation ratio for fast inference at 80B scale
Alibaba Β· 80B Β· Apache 2.0
Ultra-efficient agentic coding MoE optimized for tool-calling coding agents
Tencent Β· 80B Β· Tencent Hunyuan Community
Largest open image MoE β 13B active, strong long-prompt generation
Z.ai Β· 106B Β· MIT
Consumer-friendly GLM MoE β 12B active, strong agentic & tool use
Alibaba Β· 122B Β· Apache 2.0
Large multimodal MoE
DeepSeek Β· 158B Β· MIT
Efficient long-context V4 β 13B active, 1M context
Alibaba Β· 235B Β· Apache 2.0
Massive MoE with 22B active β frontier quality
Z.ai Β· 357B Β· MIT
Large GLM MoE with strong coding and 200K context
Alibaba Β· 397B Β· Apache 2.0
Largest multimodal Qwen 3.5 MoE
Meta Β· 400B Β· Llama 4 Community
Multimodal MoE with 128 experts β 17B active, 1M context
Alibaba Β· 480B Β· Apache 2.0
Largest open coding MoE β 35B active
DeepSeek Β· 671B Β· MIT
Massive MoE reasoning model β 37B active
DeepSeek Β· 685B Β· MIT
State-of-the-art MoE β 37B active params
Zhipu AI Β· 744B Β· MIT
MoE with 256 experts, 40B active β frontier-class agentic coding
Z.ai Β· 753B Β· MIT
Frontier open-weight coder β top SWE-bench, 1M context
Zhipu AI Β· 754B Β· MIT
Improved agentic coding β SOTA SWE-bench Pro, long-horizon tasks
Moonshot AI Β· 1.06T Β· Kimi
Natively multimodal 1T MoE β 32B active, frontier agentic
No models found
Try adjusting your search or filters
104 models Β· 284 devices Β· 22 labs
We detect your GPU or unified memory and rank open models you can run locally β chat, coding, image and video. A 9B model fits an RTX 4060; an M4 Max can take larger ones. See the scoring method, docs or install runai.
01
Open LocalAI
Visit the homepage in a desktop browser. No account or extension is required.
02
Detect or pick your hardware
Let the page read your GPU and memory, or browse the device list if detection misses your chip.
03
Compare recommended models
Review the top open models for text, coding, image and video that fit the detected VRAM or unified memory.
04
Run a model locally
Open a model page and launch it with runai, Ollama or LM Studio using the suggested quantization.
Figures assume Q4_K_M plus a small runtime overhead. Mixture-of-experts models still load all experts into memory, even when only a few are active per token.
| Memory | What usually fits | Example device |
|---|---|---|
| 8 GB | 3Bβ9B chat models at Q4, plus tiny image models. | RTX 4060 |
| 12 GB | 12Bβ14B dense models, or smaller mixture-of-experts. | RTX 4070 |
| 16 GB | 20Bβ27B at Q4, comfortable 14B at higher quality. | Apple M4 |
| 24 GB | 30B dense models, mid-size MoE, or local video. | RTX 4090 |
| 32 GB+ | Larger open MoE and high-quality image generation. | RTX 5090 |
A short mix of chat, image and video models. The list above still ranks whatever fits the hardware detected on this page.
If you want the vocabulary or the scoring math, these two pages cover it.
A 7Bβ9B chat model usually fits in 8 GB of VRAM at Q4 quantization. 12β16 GB covers most 12Bβ27B models. 24 GB and up opens 30B dense models, mid-size mixture-of-experts, and local image or video generation. On Apple Silicon the limit is unified memory, not a separate GPU frame buffer.
Local models run on your machine. Prompts never leave the device, there is no subscription meter, and you can keep working offline. Cloud assistants are usually stronger on the hardest tasks, but they need a network and send your data to a provider. LocalAI is for the open-weight models you can download and run yourself.
Quantization stores model weights in fewer bits so the file is smaller and uses less memory. Q4_K_M is the usual sweet spot for local chat: much smaller than full precision, with only a modest quality drop. Higher formats such as Q8 or F16 look closer to the original model but need more VRAM.
The site reads GPU, RAM and CPU hints from the browser, matches them against a hardware database, then scores each open model on estimated speed, memory headroom and quality. You can also skip detection and open a device page for a specific GPU, Apple chip, phone or Raspberry Pi.
Open any model page and copy its runai command, or use Ollama or LM Studio with the same GGUF weights. runai picks a quantization for your machine, downloads the file, and starts a local chat with llama.cpp.