Shortlist snapshot
| Model | Architecture / size | License | OWM hardware class | Provider origin |
|---|---|---|---|---|
| Qwen3 8BQwen / Alibaba | 8B · local general-purpose model | Apache 2.0 | ≤8B · consumer/local | China |
| Qwen3 14BQwen / Alibaba | 14B · dense general-purpose model | Apache 2.0 | 9–16B · workstation/local | China |
| Phi-4Microsoft | 14B · dense | MIT | 9–16B · workstation/local | United States |
| Granite 4.2 8BIBM | 8B · 128K enterprise reasoning | Apache 2.0 | ≤8B · consumer/local | United States |
| SmolLM3 3BHugging Face | 3B · hybrid reasoning | Apache 2.0 | ≤8B · consumer/local | United States |
| Ministral 3 14B Instruct 2512Mistral AI | 14B-class · vision-language edge model | Apache 2.0 | 9–16B · workstation/local | EU provider |
| gpt-oss-20bOpenAI | 20B · compact reasoning | Apache 2.0 | 17–32B · high-memory workstation | United States |
What should drive the decision?
A useful first-pass estimate for weight-only memory is parameters × bytes per parameter. Four-bit storage is roughly 0.5 bytes per parameter before metadata and runtime overhead; eight-bit is roughly 1 byte. Real implementations can differ materially.
Models in the 8B and 14B class generally leave more room for long context and concurrent requests than a 20B–32B model squeezed into 24 GB through aggressive quantization. That headroom can matter more than a larger checkpoint for interactive use.
If a model only barely loads, test the longest realistic prompt and generation length before calling the configuration viable. Context growth and KV cache often expose memory limits that a short smoke test misses.
Models to evaluate
Qwen3 8B
Reasoning, multilingual chat, coding and local RAG
8B-class Apache-licensed candidate with substantial memory headroom for local text work.
Qwen3 14B
Reasoning, multilingual work, coding and self-hosted assistants
14B dense candidate; quantization can preserve room for context and runtime overhead.
Phi-4
Reasoning, mathematics and code
14B MIT-licensed dense model for reasoning, mathematics and code.
Granite 4.2 8B
Enterprise assistants, reasoning, tool use and RAG
8B enterprise-oriented Apache candidate for assistants, tools and RAG.
SmolLM3 3B
Compact reasoning, agents and local inference
Compact 3B option where low memory and local iteration matter most.
Ministral 3 14B Instruct 2512
Higher-capability local chat, vision and instruction following
14B-class EU-provider vision-language option; account for multimodal components.
gpt-oss-20b
Reasoning, local and edge-class inference
Sparse reasoning candidate; validate the exact quantized/runtime configuration rather than relying on total parameters alone.
Hardware and runtime reality
Weight-only estimates are a starting point. Add KV cache, runtime workspaces, multimodal components, batching and concurrency headroom. For local inference, validate the exact quantized artifact. For server inference, measure time to first token, throughput and peak memory at target concurrency.
Long context can make an otherwise comfortable model exceed the practical memory budget. Test the longest realistic prompt and generation, not only a short loading test.
License, provider origin and Europe
Review the exact checkpoint license and any separate usage terms. Provider origin is supply-chain metadata, not an inference-location claim. If EU/EEA residency matters, map inference, RAG, embeddings, logs, telemetry, backups and subprocessors.