Runs on hardware you control
A useful local model must fit your actual computer at an acceptable speed. Downloadable weights alone do not make a model practical.
AIUpdateWatch Daily Intelligence
A comprehensive tracked list of major official general-purpose, multimodal and reasoning models up to 32B that can realistically run on consumer or workstation hardware after suitable quantization.
A useful local model must fit your actual computer at an acceptable speed. Downloadable weights alone do not make a model practical.
This table covers major official chat, instruct, multimodal and reasoning checkpoints up to 32B. It excludes unofficial fine-tunes, base-only checkpoints, embeddings, image generators, audio-only models and data-center models.
Choose by hardware first
Tiny assistants, classification and basic text work.
4 models in tableEveryday local chat, summaries and smaller multimodal tasks.
10 models in tableThe strongest practical range for many consumer systems.
8 models in tableStronger reasoning and multimodal work with more capable hardware.
6 models in tableLarge quantized models for high-memory PCs and Macs.
5 models in tableMemory bands and comfortable-hardware guidance are planning estimates, not vendor guarantees. They assume a sensible 4-bit or 5-bit quantized build, short-to-moderate context and room for the operating system. Long context, vision input, larger batches and GPU offload can require substantially more memory.
Complete tracked list
Rows are ordered from the lightest deployment tier to the most demanding. Every model name opens the provider’s official model card or documentation. The final column describes hardware for comfortable everyday use rather than the absolute minimum needed to load a model.
| Tier | Provider | Model | Parameters | Context | Input | Practical memory | Best for | License | Comfortable hardware |
|---|---|---|---|---|---|---|---|---|---|
| Very light | Gemma 3 270M — open official source in a new tab | 0.27B | 32K | Text | 4 GB+ | Classification, extraction and tiny assistants | Gemma terms | Modern 4-core CPU or integrated GPU; 16 GB system or unified memory preferred. | |
| Very light | Qwen | Qwen3-0.6B — open official source in a new tab | 0.6B | 32K | Text | 4 GB+ | Light multilingual chat and agent experiments | Apache 2.0 | Modern 4-core CPU or integrated GPU; 16 GB system or unified memory preferred. |
| Very light | Gemma 3 1B — open official source in a new tab | 1B | 32K | Text | 4–8 GB | Light chat, summaries and rewriting | Gemma terms | Modern 4-core CPU or integrated GPU; 16 GB system or unified memory preferred. | |
| Very light | Meta | Llama 3.2 1B Instruct — open official source in a new tab | 1B | 128K | Text | 4–8 GB | Mobile, edge and basic assistant tasks | Llama 3.2 | Modern 4-core CPU or integrated GPU; 16 GB system or unified memory preferred. |
| Light | DeepSeek | DeepSeek-R1-Distill-Qwen-1.5B — open official source in a new tab | 1.5B | 32K recommended | Text · reasoning | 8 GB+ | Small local math and reasoning experiments | MIT | Modern 6–8 core CPU or 6–8 GB VRAM; 16 GB system or unified memory. |
| Light | Qwen | Qwen3-1.7B — open official source in a new tab | 1.7B | 32K · 131K with YaRN | Text | 8 GB+ | Multilingual chat and efficient reasoning | Apache 2.0 | Modern 6–8 core CPU or 6–8 GB VRAM; 16 GB system or unified memory. |
| Light | IBM | Granite 3.3 2B Instruct — open official source in a new tab | 2B | 128K | Text | 8 GB+ | RAG, enterprise text and structured answers | Apache 2.0 | Modern 6–8 core CPU or 6–8 GB VRAM; 16 GB system or unified memory. |
| Light | Meta | Llama 3.2 3B Instruct — open official source in a new tab | 3B | 128K | Text | 8 GB+ | General local chat and document assistance | Llama 3.2 | Modern 6–8 core CPU or 6–8 GB VRAM; 16 GB system or unified memory. |
| Light | Hugging Face | SmolLM3-3B — open official source in a new tab | 3B | 64K · 128K with YaRN | Text | 8 GB+ | Compact chat, tools and local application use | Apache 2.0 | Modern 6–8 core CPU or 6–8 GB VRAM; 16 GB system or unified memory. |
| Light | Mistral AI | Ministral 3 3B Instruct 2512 — open official source in a new tab | 3.4B + vision encoder | 262K maximum | Text + image | 8 GB FP8 · less quantized | Multimodal edge and private local assistants | Apache 2.0 | Modern 6–8 core CPU or 6–8 GB VRAM; 16 GB system or unified memory. Keep extra memory headroom for multimodal input. |
| Light | Microsoft | Phi-4-mini-instruct — open official source in a new tab | 3.8B | 128K | Text | 8–12 GB | Reasoning, multilingual use and function calling | MIT | Modern 6–8 core CPU or 6–8 GB VRAM; 16 GB system or unified memory. |
| Light | Qwen | Qwen3-4B — open official source in a new tab | 4B | 32K · 131K with YaRN | Text | 8–12 GB | Balanced multilingual chat and reasoning | Apache 2.0 | Modern 6–8 core CPU or 6–8 GB VRAM; 16 GB system or unified memory. |
| Light | Gemma 3 4B — open official source in a new tab | 4B | 128K | Text + image | 8–12 GB | Compact multimodal assistance | Gemma terms | Modern 6–8 core CPU or 6–8 GB VRAM; 16 GB system or unified memory. Keep extra memory headroom for multimodal input. | |
| Light | Gemma 3n E2B Instruct — open official source in a new tab | 6B total · E2B effective | 32K | Text + image + audio + video | 8 GB+ with offload | Low-resource multimodal and mobile-device work | Gemma terms | Modern 6–8 core CPU or 6–8 GB VRAM; 16 GB system or unified memory. Keep extra memory headroom for multimodal input. | |
| Mainstream | Gemma 3n E4B Instruct — open official source in a new tab | 8B total · E4B effective | 32K | Text + image + audio + video | 12 GB+ with offload | Broader local multimodal understanding | Gemma terms | Modern 8-core CPU with 8–16 GB VRAM, or 24–32 GB unified memory. Keep extra memory headroom for multimodal input. | |
| Mainstream | DeepSeek | DeepSeek-R1-Distill-Qwen-7B — open official source in a new tab | 7B | 32K recommended | Text · reasoning | 12–16 GB | Local math, coding and step-by-step reasoning | MIT | Modern 8-core CPU with 8–16 GB VRAM, or 24–32 GB unified memory. |
| Mainstream | DeepSeek | DeepSeek-R1-Distill-Llama-8B — open official source in a new tab | 8B | 32K recommended | Text · reasoning | 12–16 GB | Reasoning with a Llama-based checkpoint | MIT | Modern 8-core CPU with 8–16 GB VRAM, or 24–32 GB unified memory. |
| Mainstream | Qwen | Qwen3-8B — open official source in a new tab | 8.2B | 32K · 131K with YaRN | Text | 12–16 GB | Strong multilingual chat, tools and reasoning | Apache 2.0 | Modern 8-core CPU with 8–16 GB VRAM, or 24–32 GB unified memory. |
| Mainstream | IBM | Granite 3.3 8B Instruct — open official source in a new tab | 8B | 128K | Text | 12–16 GB | Enterprise RAG, coding and structured reasoning | Apache 2.0 | Modern 8-core CPU with 8–16 GB VRAM, or 24–32 GB unified memory. |
| Mainstream | Meta | Llama 3.1 8B Instruct — open official source in a new tab | 8B | 128K | Text | 12–16 GB | Established general-purpose local assistant use | Llama 3.1 | Modern 8-core CPU with 8–16 GB VRAM, or 24–32 GB unified memory. |
| Mainstream | Mistral AI | Ministral 3 8B Instruct 2512 — open official source in a new tab | 8.4B + vision encoder | 262K maximum | Text + image | 12 GB FP8 · less quantized | Stronger multimodal edge deployment | Apache 2.0 | Modern 8-core CPU with 8–16 GB VRAM, or 24–32 GB unified memory. Keep extra memory headroom for multimodal input. |
| Mainstream | OpenAI | gpt-oss-20b — open official source in a new tab | 21B total · 3.6B active | 128K | Text · reasoning | 16 GB native MXFP4 | Local reasoning, agents, tools and structured output | Apache 2.0 | Modern 8-core CPU with 8–16 GB VRAM, or 24–32 GB unified memory. |
| Workstation | Gemma 3 12B — open official source in a new tab | 12B | 128K | Text + image | 20–24 GB | Higher-quality multimodal local work | Gemma terms | 16–24 GB VRAM, or 32–48 GB unified memory; fast SSD and strong cooling recommended. Keep extra memory headroom for multimodal input. | |
| Workstation | Mistral AI | Mistral Nemo 12B Instruct — open official source in a new tab | 12B | 128K | Text | 20–24 GB | Multilingual chat, code and function calling | Apache 2.0 | 16–24 GB VRAM, or 32–48 GB unified memory; fast SSD and strong cooling recommended. |
| Workstation | Microsoft | Phi-4 — open official source in a new tab | 14B | 16K | Text | 20–24 GB | Math, logic and dense reasoning tasks | MIT | 16–24 GB VRAM, or 32–48 GB unified memory; fast SSD and strong cooling recommended. |
| Workstation | Qwen | Qwen3-14B — open official source in a new tab | 14B | 32K · 131K with YaRN | Text | 20–24 GB | Stronger multilingual reasoning and agents | Apache 2.0 | 16–24 GB VRAM, or 32–48 GB unified memory; fast SSD and strong cooling recommended. |
| Workstation | DeepSeek | DeepSeek-R1-Distill-Qwen-14B — open official source in a new tab | 14B | 32K recommended | Text · reasoning | 20–24 GB | Stronger local reasoning, math and coding | MIT | 16–24 GB VRAM, or 32–48 GB unified memory; fast SSD and strong cooling recommended. |
| Workstation | Mistral AI | Ministral 3 14B Instruct 2512 — open official source in a new tab | 14B + vision encoder | 262K maximum | Text + image | 20–24 GB quantized | High-quality multimodal workstation use | Apache 2.0 | 16–24 GB VRAM, or 32–48 GB unified memory; fast SSD and strong cooling recommended. Keep extra memory headroom for multimodal input. |
| High-end local | Mistral AI | Mistral Small 3.1 24B Instruct — open official source in a new tab | 24B | 128K | Text + image | 32 GB quantized | High-quality private assistants and long documents | Apache 2.0 | 24–48 GB VRAM, or 48–64 GB unified memory; high-end workstation and possible multi-GPU use. Keep extra memory headroom for multimodal input. |
| High-end local | Gemma 3 27B — open official source in a new tab | 27B | 128K | Text + image | 32–40 GB | High-end multimodal work and coding | Gemma terms | 24–48 GB VRAM, or 48–64 GB unified memory; high-end workstation and possible multi-GPU use. Keep extra memory headroom for multimodal input. | |
| High-end local | Qwen | Qwen3-30B-A3B — open official source in a new tab | 30B total · 3B active | 32K · 131K with YaRN | Text | 32 GB+ | Efficient MoE reasoning, agents and multilingual work | Apache 2.0 | 24–48 GB VRAM, or 48–64 GB unified memory; high-end workstation and possible multi-GPU use. |
| High-end local | Qwen | Qwen3-32B — open official source in a new tab | 32B | 32K · 131K with YaRN | Text | 40 GB+ | High-end multilingual reasoning and agent tasks | Apache 2.0 | 24–48 GB VRAM, or 48–64 GB unified memory; high-end workstation and possible multi-GPU use. |
| High-end local | DeepSeek | DeepSeek-R1-Distill-Qwen-32B — open official source in a new tab | 32B | 32K recommended | Text · reasoning | 40 GB+ | The strongest practical DeepSeek R1 distill | MIT | 24–48 GB VRAM, or 48–64 GB unified memory; high-end workstation and possible multi-GPU use. |
A model that barely fits leaves too little memory for context, applications and the operating system. Choose one tier below your theoretical maximum for a smoother experience.
Published 128K or 262K context limits do not mean a consumer computer can use the full window comfortably. Begin with 4K–16K and increase only when necessary.
Quality, speed and memory use vary between GGUF, MLX, ONNX, FP8, MXFP4 and other packages. Evaluate the exact file and runtime you plan to keep.