Models to evaluate
| Model | Architecture / size | License | Hardware class | Provider |
|---|---|---|---|---|
| Qwen3 0.6BQwen / Alibaba | 0.6B · compact reasoning-capable text model | Apache 2.0 | ≤8B · consumer/local | 🇨🇳 China |
| Qwen3 1.7BQwen / Alibaba | 1.7B · compact reasoning-capable text model | Apache 2.0 | ≤8B · consumer/local | 🇨🇳 China |
| Qwen3 4BQwen / Alibaba | 4B · compact general-purpose model | Apache 2.0 | ≤8B · consumer/local | 🇨🇳 China |
| SmolLM3 3BHugging Face | 3B · hybrid reasoning | Apache 2.0 | ≤8B · consumer/local | 🇺🇸 United States |
| Phi-4 Mini InstructMicrosoft | compact · multilingual | MIT | ≤8B · consumer/local | 🇺🇸 United States |
| Gemma 3 1B ITGoogle DeepMind | 1B · compact text instruct | Gemma Terms | ≤8B · consumer/local | 🇺🇸 United States |
| Ministral 3 3B Instruct 2512Mistral AI | 3B-class · vision-language edge model | Apache 2.0 | ≤8B · consumer/local | 🇫🇷 France |
| Gemma 3n E2B ITGoogle DeepMind | E2B-class · mobile multimodal | Gemma Terms | ≤8B · consumer/local | 🇺🇸 United States |
Decision criteria
Prefer models whose quantized weights leave several gigabytes of working headroom.
Measure the longest realistic context; KV cache can erase apparent free memory.
CPU/GPU split can expand what loads, but may materially change latency.
For multimodal models, account for vision/audio components in addition to language weights.
Deployment reality
Validate the exact checkpoint, precision or quantization, runtime, context length and concurrency target. Weight memory alone does not capture KV cache, runtime workspaces, multimodal encoders or distributed-serving overhead.
OWM keeps license, provider origin and data residency separate. A provider-country label is provenance metadata; the deployer determines where inference and connected services run.