Models to evaluate
| Model | Architecture / size | License | Hardware class | Provider |
|---|---|---|---|---|
| Phi-4 Reasoning PlusMicrosoft | 14B · dense reasoning model | MIT | 9–16B · workstation/local | 🇺🇸 United States |
| Qwen3 14BQwen / Alibaba | 14B · dense general-purpose model | Apache 2.0 | 9–16B · workstation/local | 🇨🇳 China |
| Qwen3-32BQwen / Alibaba | 32B · dense | Apache 2.0 | 17–32B · high-memory workstation | 🇨🇳 China |
| gpt-oss-20bOpenAI | 20B · compact reasoning | Apache 2.0 | 17–32B · high-memory workstation | 🇺🇸 United States |
| Magistral Small 2506Mistral AI | 24B-class · reasoning | Apache 2.0 | 17–32B · high-memory workstation | 🇫🇷 France |
| DeepSeek-R1DeepSeek | reasoning model · MoE | MIT | Model-specific · large / specialized | 🇨🇳 China |
| Qwen3-235B-A22BQwen / Alibaba | 235B / 22B active · MoE | Apache 2.0 | Model-specific · large / specialized | 🇨🇳 China |
| GLM-4.5Z.ai | 355B / 32B active · MoE | MIT | Model-specific · large / specialized | 🇨🇳 China |
Decision criteria
Measure reasoning quality and total generated tokens together.
Include tool-use and structured-output reliability where relevant.
Compare fast/non-thinking modes separately when a model supports them.
Large MoE checkpoints need infrastructure evaluation independent of active-parameter counts.
Deployment reality
Validate the exact checkpoint, precision or quantization, runtime, context length and concurrency target. Weight memory alone does not capture KV cache, runtime workspaces, multimodal encoders or distributed-serving overhead.
OWM keeps license, provider origin and data residency separate. A provider-country label is provenance metadata; the deployer determines where inference and connected services run.