Modelos de IA local que puedes ejecutar

Busca en el catálogo abierto completo — clasificado para tu GPU, Mac u otro dispositivo.

GPUVRAMBWRAMNúcleos

FLUX.2 Klein 4B

hace 7m

Black Forest Labs · 4B · Apache 2.0

Fastest open FLUX.2 — sub-second text-to-image and multi-reference editing on consumer GPUs

2.5GB·32K ctx·

Wan 2.2 TI2V 5B

hace 1 año

Alibaba · 5B · Apache 2.0

Unified text/image-to-video — the local sweet spot under Apache 2.0

3.1GB·4K ctx·

Z-Image Turbo

hace 9m

Alibaba · 6B · Apache 2.0

8-step distilled image model — photorealism and bilingual text on 16GB cards

3.6GB·4K ctx·

HunyuanVideo 1.5

hace 9m

Tencent · 8.3B · Tencent Hunyuan Community

Compact cinematic video model — strong faces and motion on a single 4090

4.8GB·4K ctx·

Qwen3-VL 8B

hace 11m

Alibaba · 8.8B · Apache 2.0

The community-favourite local VLM — superb OCR, receipts & captioning

5GB·256K ctx·

Qwen 3.5 9B

hace 6m

Alibaba · 9B · Apache 2.0

Multimodal Qwen 3.5 mid-size

22 AA·5.1GB·32K ctx·

LTX 2.3

hace 5m

Lightricks · 19B · LTX-2 Community

Open 4K video with native stereo audio — text, image and video-to-video

10.2GB·4K ctx·

Qwen Image 2512

hace 8m

Alibaba · 20B · Apache 2.0

Open text-to-image with strong English and Chinese typography

10.7GB·4K ctx·

GPT-OSS 20B

hace 1 año

OpenAI · 21B · Apache 2.0

OpenAI's open-weight MoE with configurable reasoning

15 AA·11.3GB·128K ctx·

Gemma 4 26B-A4B IT

hace 4m

Google · 27B · Gemma

Gemma 4 MoE instruct model (official)

26 AA·14.3GB·256K ctx·

Qwen 3.8 27B

hace 26d

Alibaba · 27B · Apache 2.0

Flagship dense Qwen 3.8 — native multimodal all-rounder with video understanding

52 AA·14.3GB·256K ctx·

Wan 2.2 T2V A14B

hace 1 año

Alibaba · 27B · Apache 2.0

Flagship open Wan 2.2 — 14B-active MoE for photoreal text-to-video

14.3GB·4K ctx·

Muse Glimmer 30B

hace 26d

Meta · 30B · Apache 2.0

Open agentic 30B distilled from Muse Spark — tool use, vision and local recovery on a single GPU

35 AA·15.9GB·128K ctx·

Qwen3-VL 30B-A3B

hace 1 año

Alibaba · 31B · Apache 2.0

Efficient vision MoE — 3B active, strong temporal & document understanding

16.4GB·256K ctx·

FLUX.2 Dev

hace 9m

Black Forest Labs · 32B · FLUX Non-Commercial

Flagship open-weight FLUX.2 — text-to-image and multi-reference editing up to 4MP

16.9GB·32K ctx·

MiniMax H3

hace 26d

MiniMax · 33B · MiniMax Community

Open video generation — text/image to 2K video with native stereo audio

17.4GB·32K ctx·

Agents-A1 35B-A3B

hace 2m

InternScience · 35B · Apache 2.0

Efficient multimodal agentic MoE for long-horizon search, engineering and scientific research

18.4GB·256K ctx·

Ornith 1.0 35B-A3B

hace 2m

DeepReinforce · 35B · MIT

Agentic coding MoE with a 3B active working set and self-improving training

18.4GB·256K ctx·

Qwen 3.6 35B-A3B

hace 4m

Alibaba · 36B · Apache 2.0

Big-model quality at 3B-active speed — the mid-hardware sweet spot

32 AA·18.9GB·256K ctx·

HunyuanImage 3.0 Instruct

hace 7m

Tencent · 80B · Tencent Hunyuan Community

Reasoning image model — prompt rewrite, chain-of-thought and image-to-image editing

41.5GB·4K ctx·

Llama 4 Scout 17B

hace 1 año

Meta · 109B · Llama 4 Community

MoE with 16 experts, 17B active params

10 AA·56.3GB·128K ctx·

GPT-OSS 120B

hace 1 año

OpenAI · 117B · Apache 2.0

OpenAI's flagship open-weight MoE — 52.6% SWE-bench

24 AA·60.4GB·128K ctx·

Mistral Small 4 119B

hace 5m

Mistral AI · 119B · Apache 2.0

Sparse Mistral Small 4 — 6.5B active, strong local all-rounder

20 AA·61.5GB·256K ctx·

Qwen 3 VL 235B-A22B

hace 9m

Alibaba · 235B · Apache 2.0

Flagship vision-language MoE — frontier multimodal reasoning and agentic GUI control

120.9GB·256K ctx·

Hy3

hace 1 mes

Tencent · 295B · Apache 2.0

Production-focused agentic MoE with strong coding, tool use and long-context reasoning

42 AA·151.6GB·256K ctx·

MiniMax M3

hace 2m

MiniMax · 428B · MiniMax Community

Native multimodal MoE — understands text, image and long video with 1M context

45 AA·219.7GB·1024K ctx·

GLM-5.3

hace 26d

Z.ai · 753B · MIT

Same 753B / 40B-active base as GLM-5.2 — post-training lifts coding and long-horizon agents, 1M context

386.2GB·1024K ctx·

LongCat 2.0

hace 1 mes

Meituan · 1.6T · MIT

Frontier-scale agentic and coding MoE with sparse attention and native 1M context

34 AA·820.1GB·1024K ctx·

DeepSeek V4 Pro

hace 4m

DeepSeek · 1.6T · MIT

Flagship V4 MoE — 49B active, 1M context

820.1GB·1024K ctx·

Qwen 3.8 2.4T-A95B

hace 26d

Alibaba · 2.4T · Qwen

Frontier Qwen 3.8 MoE — 95B active, 1M context

58 AA·1229.8GB·1024K ctx·

Kimi K3

hace 1 mes

Moonshot AI · 2.8T · Kimi

Frontier 2.8T multimodal MoE — 104B active, native video understanding, 1M context

60 AA·1424.5GB·1024K ctx·
Todos los modelos

Qwen 3 0.6B

hace 1 año

Alibaba · 0.6B · Apache 2.0

Ultra-light Qwen 3 model for constrained devices

0.8GB·32K ctx·

Qwen 3.5 0.8B

hace 6m

Alibaba · 0.8B · Apache 2.0

Ultra-tiny model for embedded and edge

5* AA·0.9GB·32K ctx·

Llama 3.2 1B

hace 2a

Meta · 1B · Llama 3.2 Community

Meta's smallest Llama for edge devices

1GB·128K ctx·

Gemma 3 1B

hace 1 año

Google · 1B · Gemma

Google's tiny Gemma for on-device

1GB·32K ctx·

Wan 2.1 T2V 1.3B

hace 1 año

Alibaba · 1.3B · Apache 2.0

Tiny open text-to-video — 480p clips on 8GB consumer GPUs

1.2GB·4K ctx·

Qwen 2.5 Coder 1.5B

hace 1 año

Alibaba · 1.5B · Apache 2.0

Ultra-lightweight coding model

1.3GB·32K ctx·

DeepSeek R1 1.5B

hace 1 año

DeepSeek · 1.5B · MIT

Tiny reasoning model distilled from R1

1.3GB·64K ctx·

Qwen 3 1.7B

hace 1 año

Alibaba · 1.7B · Apache 2.0

Compact multilingual Qwen 3

1.4GB·32K ctx·

Qwen 3.5 2B

hace 6m

Alibaba · 2B · Apache 2.0

Small multimodal Qwen 3.5

7* AA·1.5GB·32K ctx·

Llama 3.2 3B

hace 2a

Meta · 3B · Llama 3.2 Community

Lightweight Llama for mobile and edge

2GB·128K ctx·

SmolLM3 3B

hace 1 año

HuggingFace · 3B · Apache 2.0

Lightweight multilingual reasoning

2GB·128K ctx·

Granite 4.1 3B

hace 4m

IBM · 3B · Apache 2.0

Compact enterprise model for edge and constrained environments

2GB·128K ctx·

Ministral 3 3B

hace 8m

Mistral AI · 3B · Apache 2.0

Current-gen tiny Ministral — edge chat with 256K context

7 AA·2GB·256K ctx·

Phi-4 Mini Reasoning

hace 1 año

Microsoft · 3.8B · MIT

Lightweight reasoning model

2.4GB·16K ctx·

Gemma 3 4B

hace 1 año

Google · 4B · Gemma

Multimodal Gemma with 128K context

2.5GB·128K ctx·

Qwen 3.5 4B

hace 6m

Alibaba · 4B · Apache 2.0

Small multimodal Qwen 3.5

20* AA·2.5GB·32K ctx·

Qwen3-VL 4B

hace 11m

Alibaba · 4.4B · Apache 2.0

Compact dedicated vision-language model — OCR & image chat on edge

2.8GB·256K ctx·

Gemma 4 E2B IT

hace 4m

Google · 5B · Gemma

Gemma 4 efficient instruct model (official)

10* AA·3.1GB·256K ctx·

Qwen 2.5 Coder 7B

hace 1 año

Alibaba · 7B · Apache 2.0

Dedicated coding model

4.1GB·128K ctx·

DeepSeek R1 Distill 7B

hace 1 año

DeepSeek · 7B · MIT

R1 reasoning distilled into Qwen 7B

4.1GB·64K ctx·

Gemma 4 E4B IT

hace 4m

Google · 8B · Gemma

Gemma 4 balanced instruct model (official)

12* AA·4.6GB·256K ctx·

Llama 3.1 8B

hace 2a

Meta · 8B · Llama 3.1 Community

Meta's versatile 8B — great quality/speed ratio

4.6GB·128K ctx·

Qwen 3 8B

hace 1 año

Alibaba · 8B · Apache 2.0

Qwen 3 with thinking mode support

4.6GB·128K ctx·

Granite 4.1 8B

hace 4m

IBM · 8B · Apache 2.0

Balanced general-purpose enterprise model

4.6GB·128K ctx·

Ministral 8B

hace 1 año

Mistral AI · 8B · MRL

Mistral's efficient 8B model

4.6GB·32K ctx·

GLM-4 9B

hace 2a

Zhipu AI · 9B · GLM-4

Multilingual model supporting 26 languages with 128K context

5.1GB·128K ctx·

Nemotron Nano 9B v2

hace 1 año

NVIDIA · 9B · NVIDIA Open

Hybrid Mamba2 architecture for reasoning

9* AA·5.1GB·128K ctx·

Ornith 1.0 9B

hace 2m

DeepReinforce · 9B · MIT

Self-improving agentic coding model optimized for terminal and software engineering tasks

5.1GB·256K ctx·

FLUX.2 Klein 9B

hace 7m

Black Forest Labs · 9B · FLUX Non-Commercial

Higher-quality distilled FLUX.2 — sub-second generation and multi-reference editing

5.1GB·32K ctx·

Gemma 3 12B

hace 1 año

Google · 12B · Gemma

Multimodal Gemma with 128K context

6.6GB·128K ctx·

Mistral Nemo 12B

hace 2a

Mistral AI · 12B · Apache 2.0

Multilingual 12B with 128K context

6.6GB·128K ctx·

Gemma 4 12B IT

hace 4m

Google · 12B · Apache 2.0

Gemma 4 mid-size instruct — multimodal any-to-any

22* AA·6.6GB·256K ctx·

Phi-4 14B

hace 1 año

Microsoft · 14B · MIT

Microsoft's reasoning-focused model

5* AA·7.7GB·16K ctx·

Qwen 3 14B

hace 1 año

Alibaba · 14B · Apache 2.0

Strong all-rounder with thinking mode

7.7GB·128K ctx·

DeepSeek R1 Distill 14B

hace 1 año

DeepSeek · 14B · MIT

R1 reasoning distilled into Qwen 14B

7.7GB·64K ctx·

Ministral 3 14B

hace 8m

Mistral AI · 14B · Apache 2.0

Current-gen Ministral mid-size — local assistant with 256K context

11 AA·7.7GB·256K ctx·

LFM2 24B

hace 9m

Liquid AI · 24B · Liquid AI

Hybrid MoE with convolution+attention layers — 2.3B active

5* AA·12.8GB·32K ctx·

Devstral Small 2 24B

hace 8m

Mistral AI · 24B · Apache 2.0

Coding-focused model with 256K context — 68% SWE-bench

12.8GB·256K ctx·

Mistral Small 3.1 24B

hace 1 año

Mistral AI · 24B · Apache 2.0

Multimodal Mistral with vision support

12.8GB·128K ctx·

DiffusionGemma 26B-A4B IT

hace 2m

Google · 26B · Apache 2.0

Discrete diffusion MoE — 1100+ tok/s on H100, multimodal (text/image/video)

13* AA·13.8GB·256K ctx·

Qwen 3.5 27B

hace 6m

Alibaba · 27.8B · Apache 2.0

Flagship native multimodal Qwen 3.5

14.7GB·256K ctx·

Qwen 3.6 27B

hace 4m

Alibaba · 27.8B · Apache 2.0

Flagship dense Qwen 3.6 — native multimodal all-rounder

38 AA·14.7GB·256K ctx·

Qwen 3 30B-A3B

hace 1 año

Alibaba · 30B · Apache 2.0

MoE with only 3.3B active — extremely efficient

15.9GB·128K ctx·

Nemotron 3 Nano 30B

hace 1 año

NVIDIA · 30B · NVIDIA Open

MoE with 1M context and 3B active

15 AA·15.9GB·1024K ctx·

Granite 4.1 30B

hace 4m

IBM · 30B · Apache 2.0

High-capacity enterprise model for complex reasoning and tool use

15.9GB·128K ctx·

North Mini Code

hace 2m

Cohere · 30B · Apache 2.0

Open agentic coding MoE with 3B active — built for software engineering and terminal tasks

15.9GB·256K ctx·

Qwen 3 Coder 30B-A3B

hace 1 año

Alibaba · 30B · Apache 2.0

Efficient agentic coding MoE — 3B active, 256K context

15.9GB·256K ctx·

Qwen 3 32B

hace 1 año

Alibaba · 32B · Apache 2.0

Qwen 3 flagship dense model

16.9GB·128K ctx·

DeepSeek R1 Distill 32B

hace 1 año

DeepSeek · 32B · MIT

R1 reasoning distilled into Qwen 32B — sweet spot

16.9GB·64K ctx·

OLMo 2 32B

hace 1 año

Allen AI · 32B · Apache 2.0

Fully open research model by Allen AI

16.9GB·4K ctx·

Gemma 4 31B IT

hace 4m

Google · 33B · Gemma

Gemma 4 flagship instruct model (official)

30 AA·17.4GB·256K ctx·

Gemma 4 31B

hace 4m

Google · 33B · Gemma

Gemma 4 flagship base model (official)

17.4GB·256K ctx·

Command R 35B

hace 2a

Cohere · 35B · CC BY-NC 4.0

Optimized for retrieval-augmented generation

18.4GB·128K ctx·

Qwen 3.5 35B-A3B

hace 6m

Alibaba · 35B · Apache 2.0

Efficient multimodal MoE with 3B active

18.4GB·256K ctx·

Mixtral 8x7B

hace 2a

Mistral AI · 47B · Apache 2.0

MoE with 12.9B active params

24.6GB·32K ctx·

Llama 3.3 70B

hace 1 año

Meta · 70B · Llama 3.3 Community

Best open model at 70B class

9* AA·36.4GB·128K ctx·

Qwen 3 Next 80B-A3B

hace 8m

Alibaba · 80B · Apache 2.0

High-sparsity MoE — extreme low activation ratio for fast inference at 80B scale

41.5GB·256K ctx·

Qwen 3 Coder Next 80B-A3B

hace 6m

Alibaba · 80B · Apache 2.0

Ultra-efficient agentic coding MoE optimized for tool-calling coding agents

41.5GB·256K ctx·

HunyuanImage 3.0

hace 1 año

Tencent · 80B · Tencent Hunyuan Community

Largest open image MoE — 13B active, strong long-prompt generation

41.5GB·4K ctx·

GLM-4.5 Air

hace 1 año

Z.ai · 106B · MIT

Consumer-friendly GLM MoE — 12B active, strong agentic & tool use

54.8GB·128K ctx·

Qwen 3.5 122B-A10B

hace 6m

Alibaba · 122B · Apache 2.0

Large multimodal MoE

33 AA·63GB·256K ctx·

DeepSeek V4 Flash

hace 4m

DeepSeek · 158B · MIT

Efficient long-context V4 — 13B active, 1M context

81.4GB·1024K ctx·

Qwen 3 235B-A22B

hace 1 año

Alibaba · 235B · Apache 2.0

Massive MoE with 22B active — frontier quality

120.9GB·128K ctx·

GLM-4.6

hace 1 año

Z.ai · 357B · MIT

Large GLM MoE with strong coding and 200K context

183.4GB·195K ctx·

Qwen 3.5 397B-A17B

hace 6m

Alibaba · 397B · Apache 2.0

Largest multimodal Qwen 3.5 MoE

34 AA·203.9GB·256K ctx·

Llama 4 Maverick 17B-128E

hace 1 año

Meta · 400B · Llama 4 Community

Multimodal MoE with 128 experts — 17B active, 1M context

14 AA·205.4GB·1024K ctx·

Qwen 3 Coder 480B

hace 1 año

Alibaba · 480B · Apache 2.0

Largest open coding MoE — 35B active

246.4GB·256K ctx·

DeepSeek R1

hace 1 año

DeepSeek · 671B · MIT

Massive MoE reasoning model — 37B active

344.2GB·64K ctx·

DeepSeek V3.2

hace 8m

DeepSeek · 685B · MIT

State-of-the-art MoE — 37B active params

351.4GB·128K ctx·

GLM-5

hace 6m

Zhipu AI · 744B · MIT

MoE with 256 experts, 40B active — frontier-class agentic coding

381.6GB·128K ctx·

GLM-5.2

hace 2m

Z.ai · 753B · MIT

Frontier open-weight coder — top SWE-bench, 1M context

53 AA·386.2GB·1024K ctx·

GLM-5.1

hace 4m

Zhipu AI · 754B · MIT

Improved agentic coding — SOTA SWE-bench Pro, long-horizon tasks

386.7GB·128K ctx·

Kimi K2.6

hace 4m

Moonshot AI · 1.06T · Kimi

Natively multimodal 1T MoE — 32B active, frontier agentic

542.4GB·256K ctx·

Todos los nombres de producto, logos y marcas son propiedad de sus respectivos dueños. Apple, NVIDIA, AMD, Intel, Qualcomm y todos los nombres de modelos de IA mencionados en este sitio son marcas registradas de sus titulares. Este sitio no está afiliado ni respaldado por ninguna de estas empresas.