Hardware-oriented guide

Best Open-Weight Models for Local Deployment

Choose by memory and workload — not by leaderboard position.

This guide filters the 32-model registry for checkpoints with documented local, workstation, edge or quantized deployment paths. “Best” here means best-fit candidates by hardware class and use case, not a universal model ranking.

Start with total parametersActive MoE parameters are not checkpoint size.
Then choose precisionFP16/BF16, 8-bit and 4-bit materially change weight memory.
Reserve runtime memoryKV cache, context, batching and vision add overhead.
Check the licenseLocal control does not override usage terms.
How to use this guide

Hardware fit comes before model hype.

A local model is only useful if the weights, cache and runtime fit the hardware you actually control.

OpenWeightModels uses total parameter count as a first navigation layer, then checks the model passport for publisher-documented local runtimes, quantization paths and hardware notes. This is intentionally more conservative than inferring “runs locally” from weight availability alone.

Small dense models such as SmolLM3 3B or Phi-4 Mini can be reasonable consumer-device candidates after quantization. Mid-sized 20–32B models can be strong workstation options, but long contexts and multimodal workloads can move them into much heavier memory classes. Sparse MoE models can reduce active compute while still requiring the full expert checkpoint to be stored.

Compact local models: up to 8B

The easiest part of the registry to experiment with on consumer hardware. Quantization can make these models practical on laptops, desktops and smaller GPUs, depending on runtime and context.

SmolLM3 3B

At 3B parameters SmolLM3 is one of the most accessible full passports in the registry. Quantization makes it practical on consumer and edge-class hardware, but 128K extended context can still overwhelm small-memory devices because KV cache grows with sequence length.

3BApache 2.0MLX / MLC, ONNX, Transformers

Phi-4 Mini Instruct

At 3.8B parameters, the model is suitable for consumer GPUs and some CPU/edge deployments after quantization. The 128K window can nevertheless consume significant KV-cache memory; small weights do not mean small long-context runtime memory.

3.8BMITMicrosoft model ecosystem, ONNX / edge tooling, Transformers

Gemma 3 4B IT

Google positions the 4B class for desktops, smaller servers and quantized local use. Exact RAM/VRAM depends strongly on quantization, image token count, context length and runtime; 128K workloads require substantially more cache than short chats.

4BGemma TermsOllama, Transformers, Vertex AI Model Garden

Mathstral 7B v0.1

At 7B dense parameters, Mathstral is accessible for local inference, especially under 8-bit or 4-bit quantization. Mathematical workloads can still generate long derivations, so output length and batch sizing matter.

7BApache 2.0Transformers, mistral-inference, vLLM

Gemma 3n E4B IT

Gemma 3n is designed to lower device memory pressure through architectural offloading and conditional modality loading. Exact mobile memory use depends on active modalities, precision and runtime; “E4B” should not be interpreted as a conventional exactly-4B dense checkpoint.

~8B raw / E4B effective profileGemma TermsGoogle AI Edge ecosystem, On-device/mobile runtimes, Quantized ecosystem

Granite 3.3 8B Instruct

An 8B dense model is practical on a single modern GPU and on CPU/unified-memory systems under quantization. 128K context can become the dominant memory consumer for long sessions, so weight size alone is not sufficient sizing.

8BApache 2.0IBM watsonx ecosystem, Quantized local variants, Transformers

Workstation-friendly models: 9B to 16B

A useful middle tier for stronger local inference. Full precision can still be demanding, but quantized deployments are often practical on modern workstation hardware.

Gemma 3 12B IT

Google positions the 12B class for high-end desktops and servers. Exact RAM/VRAM depends strongly on quantization, image token count, context length and runtime; 128K workloads require substantially more cache than short chats.

12BGemma TermsOllama, Transformers, Vertex AI Model Garden

Mistral Nemo Instruct 2407

A 12B dense model is comparatively approachable: BF16 still needs substantial memory, while 8-bit/4-bit builds can fit higher-end consumer hardware. Long 128K sessions add cache overhead independent of weight size.

12BApache 2.0NVIDIA NeMo, Transformers, mistral-inference

Phi-4

14B dense parameters make Phi-4 far easier to host than frontier-scale models. BF16 remains a high-memory single-GPU or small multi-GPU target; quantization can bring it to consumer-class hardware. The 16K context limits cache growth relative to 128K models.

14BMITAzure / Microsoft model catalog, ONNX / Microsoft optimization ecosystem, Transformers

High-memory local models: 17B to 32B

These models can be locally deployable, but the word “local” now means high-memory workstation, multi-GPU or unified-memory system rather than an ordinary laptop.

gpt-oss-20b

OpenAI positioned gpt-oss-20b for memory-constrained and local scenarios and stated that it can run with roughly 16 GB of memory in suitable deployments. Real requirements vary with context, runtime and acceleration.

20.9BApache 2.0Ollama / local apps, SGLang, Transformers

Devstral Small 2505

The publisher explicitly positions Devstral as locally deployable after quantization. Repository-scale context and coding-agent tool traces can still raise memory/latency substantially compared with short completion workloads.

24BApache 2.0Local quantized deployment, OpenHands-style agent scaffold, mistral-common / mistral-inference

Magistral Small 2506

Quantized local deployment is realistic for a 24B model, but long chain-of-thought outputs and large context windows increase KV cache and generation time. Mistral explicitly recommends a 40K maximum model length for quality even though the architecture exposes 128K.

24BApache 2.0Mistral tooling, Quantized local deployment, vLLM

Mistral Small 3.1 24B Instruct

At 24B dense parameters, full precision is a serious workstation/server model but quantization can bring it into single-high-end-GPU or 32GB unified-memory territory. Vision and 128K context still increase runtime memory beyond weight storage alone.

24BApache 2.0Quantized local runtimes, Transformers, mistral-inference

Gemma 3 27B IT

Google positions the 27B class for large servers and server clusters. Exact RAM/VRAM depends strongly on quantization, image token count, context length and runtime; 128K workloads require substantially more cache than short chats.

27BGemma TermsOllama, Transformers, Vertex AI Model Garden

Qwen3-30B-A3B

The 30.5B total size makes this MoE much more accessible than giant sparse models. However, 3.3B active parameters should still not be interpreted as a 3B storage footprint; all experts remain part of the checkpoint.

30.5BApache 2.0Quantized local builds, SGLang, Transformers

Qwen2.5-Coder-32B-Instruct

A dense 32B model is feasible on multi-GPU servers and high-memory workstations with quantization. Large repository contexts and long generations increase KV-cache and latency requirements.

~32.5BApache 2.0GGUF / quantized ecosystem, SGLang, Transformers

Qwen3-32B

A 32.8B dense checkpoint is much easier to plan than a 235B MoE model because its parameter count and active compute are aligned. BF16 still requires substantial accelerator memory; 4-bit quantization moves it into high-memory workstation territory.

32.8BApache 2.0SGLang, Transformers, llama.cpp / Ollama ecosystem
Large-model exception

Local control is not the same as consumer hardware.

gpt-oss-120b is a useful boundary case.

OpenAI documents that the 120B checkpoint was designed to fit on a single 80 GB accelerator. That makes independent single-accelerator deployment possible in the right environment, but it is not a normal desktop model. The same distinction applies across the registry: “self-hosted” can mean anything from a laptop to an eight-GPU server.

For giant models such as Llama 4 Maverick, Kimi K2 or Qwen3-Coder 480B, independent deployment remains datacenter-class even when sparse activation reduces per-token compute.

FAQ

Local deployment questions.

For exact hardware and runtime details, follow the linked Model Passport rather than treating a guide tier as a minimum requirement.

What is the best open-weight model to run locally?

There is no single best model for every device. The useful choice depends on memory, modality, task, context length and license. This guide groups verified models by total parameter scale and documented local deployment paths.

How much VRAM do I need for a local AI model?

Weight precision is only one part of memory use. Context length, KV cache, batch size, runtime overhead and multimodal inputs can materially increase RAM or VRAM requirements.

Does a Mixture-of-Experts model with 3B active parameters fit like a 3B dense model?

No. Active parameters describe per-token compute, while the full expert set still contributes to checkpoint storage and memory planning.

Which licenses are simplest for local commercial deployment?

Apache 2.0 and MIT are permissive software licenses in this registry, but model-specific policies or application rules can still matter. Custom licenses should be reviewed separately.