These models can be locally deployable, but the word “local” now means high-memory workstation, multi-GPU or unified-memory system rather than an ordinary laptop.
OpenAI positioned gpt-oss-20b for memory-constrained and local scenarios and stated that it can run with roughly 16 GB of memory in suitable deployments. Real requirements vary with context, runtime and acceleration.
20.9BApache 2.0Ollama / local apps, SGLang, Transformers
The publisher explicitly positions Devstral as locally deployable after quantization. Repository-scale context and coding-agent tool traces can still raise memory/latency substantially compared with short completion workloads.
24BApache 2.0Local quantized deployment, OpenHands-style agent scaffold, mistral-common / mistral-inference
Quantized local deployment is realistic for a 24B model, but long chain-of-thought outputs and large context windows increase KV cache and generation time. Mistral explicitly recommends a 40K maximum model length for quality even though the architecture exposes 128K.
24BApache 2.0Mistral tooling, Quantized local deployment, vLLM
At 24B dense parameters, full precision is a serious workstation/server model but quantization can bring it into single-high-end-GPU or 32GB unified-memory territory. Vision and 128K context still increase runtime memory beyond weight storage alone.
24BApache 2.0Quantized local runtimes, Transformers, mistral-inference
Google positions the 27B class for large servers and server clusters. Exact RAM/VRAM depends strongly on quantization, image token count, context length and runtime; 128K workloads require substantially more cache than short chats.
27BGemma TermsOllama, Transformers, Vertex AI Model Garden
The 30.5B total size makes this MoE much more accessible than giant sparse models. However, 3.3B active parameters should still not be interpreted as a 3B storage footprint; all experts remain part of the checkpoint.
30.5BApache 2.0Quantized local builds, SGLang, Transformers
A dense 32B model is feasible on multi-GPU servers and high-memory workstations with quantization. Large repository contexts and long generations increase KV-cache and latency requirements.
~32.5BApache 2.0GGUF / quantized ecosystem, SGLang, Transformers
A 32.8B dense checkpoint is much easier to plan than a 235B MoE model because its parameter count and active compute are aligned. BF16 still requires substantial accelerator memory; 4-bit quantization moves it into high-memory workstation territory.
32.8BApache 2.0SGLang, Transformers, llama.cpp / Ollama ecosystem