Models to evaluate
| Model | Architecture / size | License | Hardware class | Provider |
|---|---|---|---|---|
| Qwen3 8BQwen / Alibaba | 8B · local general-purpose model | Apache 2.0 | ≤8B · consumer/local | 🇨🇳 China |
| Qwen3 14BQwen / Alibaba | 14B · dense general-purpose model | Apache 2.0 | 9–16B · workstation/local | 🇨🇳 China |
| Phi-4Microsoft | 14B · dense | MIT | 9–16B · workstation/local | 🇺🇸 United States |
| Gemma 3 12B ITGoogle DeepMind | 12B · multimodal | Gemma Terms | 9–16B · workstation/local | 🇺🇸 United States |
| Mistral Nemo Instruct 2407Mistral AI / NVIDIA | 12B-class · multilingual | Apache 2.0 | 9–16B · workstation/local | 🇫🇷 France |
| Ministral 3 8B Instruct 2512Mistral AI | 8B-class · vision-language edge model | Apache 2.0 | ≤8B · consumer/local | 🇫🇷 France |
| Granite 4.2 8BIBM | 8B · 128K enterprise reasoning | Apache 2.0 | ≤8B · consumer/local | 🇺🇸 United States |
| Qwen2.5-Coder 7B InstructQwen / Alibaba | 7B · code-specialized instruct | Apache 2.0 | ≤8B · consumer/local | 🇨🇳 China |
Decision criteria
8B-class models usually provide the most comfortable local headroom.
12B–14B models often benefit from 4-bit or other memory-efficient formats.
Long-context work can require more conservative model sizing.
Evaluate quality per watt and latency, not only the largest model that fits.
Deployment reality
Validate the exact checkpoint, precision or quantization, runtime, context length and concurrency target. Weight memory alone does not capture KV cache, runtime workspaces, multimodal encoders or distributed-serving overhead.
OWM keeps license, provider origin and data residency separate. A provider-country label is provenance metadata; the deployer determines where inference and connected services run.