Models to evaluate
| Model | Architecture / size | License | Hardware class | Provider |
|---|---|---|---|---|
| Granite 4.2 8BIBM | 8B · 128K enterprise reasoning | Apache 2.0 | ≤8B · consumer/local | 🇺🇸 United States |
| Qwen3 8BQwen / Alibaba | 8B · local general-purpose model | Apache 2.0 | ≤8B · consumer/local | 🇨🇳 China |
| Qwen3 14BQwen / Alibaba | 14B · dense general-purpose model | Apache 2.0 | 9–16B · workstation/local | 🇨🇳 China |
| Qwen3-32BQwen / Alibaba | 32B · dense | Apache 2.0 | 17–32B · high-memory workstation | 🇨🇳 China |
| Mistral Nemo Instruct 2407Mistral AI / NVIDIA | 12B-class · multilingual | Apache 2.0 | 9–16B · workstation/local | 🇫🇷 France |
| Ministral 3 8B Instruct 2512Mistral AI | 8B-class · vision-language edge model | Apache 2.0 | ≤8B · consumer/local | 🇫🇷 France |
| gpt-oss-20bOpenAI | 20B · compact reasoning | Apache 2.0 | 17–32B · high-memory workstation | 🇺🇸 United States |
| OLMo 3 32BAi2 | 32B · fully open research stack | Apache 2.0 | 17–32B · high-memory workstation | 🇺🇸 United States |
Decision criteria
Use the same retrieval corpus and retriever when comparing generators.
Score citation/grounding behavior separately from fluency.
Test abstention when the answer is not present in retrieved evidence.
Map embeddings, vector stores, logs and backups as part of the data path.
Deployment reality
Validate the exact checkpoint, precision or quantization, runtime, context length and concurrency target. Weight memory alone does not capture KV cache, runtime workspaces, multimodal encoders or distributed-serving overhead.
OWM keeps license, provider origin and data residency separate. A provider-country label is provenance metadata; the deployer determines where inference and connected services run.