AIUpdateWatch Daily Intelligence

Local AI models: complete practical comparison

A comprehensive tracked list of major official general-purpose, multimodal and reasoning models up to 32B that can realistically run on consumer or workstation hardware after suitable quantization.

Morning edition: Friday, July 24, 2026 · Data cutoff Jul 23, 2026, 8:30 PM (America/New_York)
Local model

Runs on hardware you control

A useful local model must fit your actual computer at an acceptable speed. Downloadable weights alone do not make a model practical.

Comparison scope

33 official models

This table covers major official chat, instruct, multimodal and reasoning checkpoints up to 32B. It excludes unofficial fine-tunes, base-only checkpoints, embeddings, image generators, audio-only models and data-center models.

Choose by hardware first

Practical local deployment tiers

Compare local hardware →
Very light4–8 GB

Tiny assistants, classification and basic text work.

4 models in table
Light8–12 GB

Everyday local chat, summaries and smaller multimodal tasks.

10 models in table
Mainstream12–16 GB

The strongest practical range for many consumer systems.

8 models in table
Workstation20–24 GB

Stronger reasoning and multimodal work with more capable hardware.

6 models in table
High-end local32–40 GB+

Large quantized models for high-memory PCs and Macs.

5 models in table

Memory bands and comfortable-hardware guidance are planning estimates, not vendor guarantees. They assume a sensible 4-bit or 5-bit quantized build, short-to-moderate context and room for the operating system. Long context, vision input, larger batches and GPU offload can require substantially more memory.

Complete tracked list

Local model comparison table

33 models

Rows are ordered from the lightest deployment tier to the most demanding. Every model name opens the provider’s official model card or documentation. The final column describes hardware for comfortable everyday use rather than the absolute minimum needed to load a model.

Comparison of major official local-capable AI models
TierProviderModelParametersContextInputPractical memoryBest forLicenseComfortable hardware
Very lightGoogleGemma 3 270M — open official source in a new tab0.27B32KText4 GB+Classification, extraction and tiny assistantsGemma termsModern 4-core CPU or integrated GPU; 16 GB system or unified memory preferred.
Very lightQwenQwen3-0.6B — open official source in a new tab0.6B32KText4 GB+Light multilingual chat and agent experimentsApache 2.0Modern 4-core CPU or integrated GPU; 16 GB system or unified memory preferred.
Very lightGoogleGemma 3 1B — open official source in a new tab1B32KText4–8 GBLight chat, summaries and rewritingGemma termsModern 4-core CPU or integrated GPU; 16 GB system or unified memory preferred.
Very lightMetaLlama 3.2 1B Instruct — open official source in a new tab1B128KText4–8 GBMobile, edge and basic assistant tasksLlama 3.2Modern 4-core CPU or integrated GPU; 16 GB system or unified memory preferred.
LightDeepSeekDeepSeek-R1-Distill-Qwen-1.5B — open official source in a new tab1.5B32K recommendedText · reasoning8 GB+Small local math and reasoning experimentsMITModern 6–8 core CPU or 6–8 GB VRAM; 16 GB system or unified memory.
LightQwenQwen3-1.7B — open official source in a new tab1.7B32K · 131K with YaRNText8 GB+Multilingual chat and efficient reasoningApache 2.0Modern 6–8 core CPU or 6–8 GB VRAM; 16 GB system or unified memory.
LightIBMGranite 3.3 2B Instruct — open official source in a new tab2B128KText8 GB+RAG, enterprise text and structured answersApache 2.0Modern 6–8 core CPU or 6–8 GB VRAM; 16 GB system or unified memory.
LightMetaLlama 3.2 3B Instruct — open official source in a new tab3B128KText8 GB+General local chat and document assistanceLlama 3.2Modern 6–8 core CPU or 6–8 GB VRAM; 16 GB system or unified memory.
LightHugging FaceSmolLM3-3B — open official source in a new tab3B64K · 128K with YaRNText8 GB+Compact chat, tools and local application useApache 2.0Modern 6–8 core CPU or 6–8 GB VRAM; 16 GB system or unified memory.
LightMistral AIMinistral 3 3B Instruct 2512 — open official source in a new tab3.4B + vision encoder262K maximumText + image8 GB FP8 · less quantizedMultimodal edge and private local assistantsApache 2.0Modern 6–8 core CPU or 6–8 GB VRAM; 16 GB system or unified memory. Keep extra memory headroom for multimodal input.
LightMicrosoftPhi-4-mini-instruct — open official source in a new tab3.8B128KText8–12 GBReasoning, multilingual use and function callingMITModern 6–8 core CPU or 6–8 GB VRAM; 16 GB system or unified memory.
LightQwenQwen3-4B — open official source in a new tab4B32K · 131K with YaRNText8–12 GBBalanced multilingual chat and reasoningApache 2.0Modern 6–8 core CPU or 6–8 GB VRAM; 16 GB system or unified memory.
LightGoogleGemma 3 4B — open official source in a new tab4B128KText + image8–12 GBCompact multimodal assistanceGemma termsModern 6–8 core CPU or 6–8 GB VRAM; 16 GB system or unified memory. Keep extra memory headroom for multimodal input.
LightGoogleGemma 3n E2B Instruct — open official source in a new tab6B total · E2B effective32KText + image + audio + video8 GB+ with offloadLow-resource multimodal and mobile-device workGemma termsModern 6–8 core CPU or 6–8 GB VRAM; 16 GB system or unified memory. Keep extra memory headroom for multimodal input.
MainstreamGoogleGemma 3n E4B Instruct — open official source in a new tab8B total · E4B effective32KText + image + audio + video12 GB+ with offloadBroader local multimodal understandingGemma termsModern 8-core CPU with 8–16 GB VRAM, or 24–32 GB unified memory. Keep extra memory headroom for multimodal input.
MainstreamDeepSeekDeepSeek-R1-Distill-Qwen-7B — open official source in a new tab7B32K recommendedText · reasoning12–16 GBLocal math, coding and step-by-step reasoningMITModern 8-core CPU with 8–16 GB VRAM, or 24–32 GB unified memory.
MainstreamDeepSeekDeepSeek-R1-Distill-Llama-8B — open official source in a new tab8B32K recommendedText · reasoning12–16 GBReasoning with a Llama-based checkpointMITModern 8-core CPU with 8–16 GB VRAM, or 24–32 GB unified memory.
MainstreamQwenQwen3-8B — open official source in a new tab8.2B32K · 131K with YaRNText12–16 GBStrong multilingual chat, tools and reasoningApache 2.0Modern 8-core CPU with 8–16 GB VRAM, or 24–32 GB unified memory.
MainstreamIBMGranite 3.3 8B Instruct — open official source in a new tab8B128KText12–16 GBEnterprise RAG, coding and structured reasoningApache 2.0Modern 8-core CPU with 8–16 GB VRAM, or 24–32 GB unified memory.
MainstreamMetaLlama 3.1 8B Instruct — open official source in a new tab8B128KText12–16 GBEstablished general-purpose local assistant useLlama 3.1Modern 8-core CPU with 8–16 GB VRAM, or 24–32 GB unified memory.
MainstreamMistral AIMinistral 3 8B Instruct 2512 — open official source in a new tab8.4B + vision encoder262K maximumText + image12 GB FP8 · less quantizedStronger multimodal edge deploymentApache 2.0Modern 8-core CPU with 8–16 GB VRAM, or 24–32 GB unified memory. Keep extra memory headroom for multimodal input.
MainstreamOpenAIgpt-oss-20b — open official source in a new tab21B total · 3.6B active128KText · reasoning16 GB native MXFP4Local reasoning, agents, tools and structured outputApache 2.0Modern 8-core CPU with 8–16 GB VRAM, or 24–32 GB unified memory.
WorkstationGoogleGemma 3 12B — open official source in a new tab12B128KText + image20–24 GBHigher-quality multimodal local workGemma terms16–24 GB VRAM, or 32–48 GB unified memory; fast SSD and strong cooling recommended. Keep extra memory headroom for multimodal input.
WorkstationMistral AIMistral Nemo 12B Instruct — open official source in a new tab12B128KText20–24 GBMultilingual chat, code and function callingApache 2.016–24 GB VRAM, or 32–48 GB unified memory; fast SSD and strong cooling recommended.
WorkstationMicrosoftPhi-4 — open official source in a new tab14B16KText20–24 GBMath, logic and dense reasoning tasksMIT16–24 GB VRAM, or 32–48 GB unified memory; fast SSD and strong cooling recommended.
WorkstationQwenQwen3-14B — open official source in a new tab14B32K · 131K with YaRNText20–24 GBStronger multilingual reasoning and agentsApache 2.016–24 GB VRAM, or 32–48 GB unified memory; fast SSD and strong cooling recommended.
WorkstationDeepSeekDeepSeek-R1-Distill-Qwen-14B — open official source in a new tab14B32K recommendedText · reasoning20–24 GBStronger local reasoning, math and codingMIT16–24 GB VRAM, or 32–48 GB unified memory; fast SSD and strong cooling recommended.
WorkstationMistral AIMinistral 3 14B Instruct 2512 — open official source in a new tab14B + vision encoder262K maximumText + image20–24 GB quantizedHigh-quality multimodal workstation useApache 2.016–24 GB VRAM, or 32–48 GB unified memory; fast SSD and strong cooling recommended. Keep extra memory headroom for multimodal input.
High-end localMistral AIMistral Small 3.1 24B Instruct — open official source in a new tab24B128KText + image32 GB quantizedHigh-quality private assistants and long documentsApache 2.024–48 GB VRAM, or 48–64 GB unified memory; high-end workstation and possible multi-GPU use. Keep extra memory headroom for multimodal input.
High-end localGoogleGemma 3 27B — open official source in a new tab27B128KText + image32–40 GBHigh-end multimodal work and codingGemma terms24–48 GB VRAM, or 48–64 GB unified memory; high-end workstation and possible multi-GPU use. Keep extra memory headroom for multimodal input.
High-end localQwenQwen3-30B-A3B — open official source in a new tab30B total · 3B active32K · 131K with YaRNText32 GB+Efficient MoE reasoning, agents and multilingual workApache 2.024–48 GB VRAM, or 48–64 GB unified memory; high-end workstation and possible multi-GPU use.
High-end localQwenQwen3-32B — open official source in a new tab32B32K · 131K with YaRNText40 GB+High-end multilingual reasoning and agent tasksApache 2.024–48 GB VRAM, or 48–64 GB unified memory; high-end workstation and possible multi-GPU use.
High-end localDeepSeekDeepSeek-R1-Distill-Qwen-32B — open official source in a new tab32B32K recommendedText · reasoning40 GB+The strongest practical DeepSeek R1 distillMIT24–48 GB VRAM, or 48–64 GB unified memory; high-end workstation and possible multi-GPU use.

How to choose from the table

Start below your limit

A model that barely fits leaves too little memory for context, applications and the operating system. Choose one tier below your theoretical maximum for a smoother experience.

Do not chase context numbers

Published 128K or 262K context limits do not mean a consumer computer can use the full window comfortably. Begin with 4K–16K and increase only when necessary.

Test the exact quantization

Quality, speed and memory use vary between GGUF, MLX, ONNX, FP8, MXFP4 and other packages. Evaluate the exact file and runtime you plan to keep.