Models to evaluate
| Model | Architecture / size | License | Hardware class | Provider |
|---|---|---|---|---|
| Qwen3-Coder-30B-A3B-InstructQwen / Alibaba | 30B / ~3B active · coding MoE | Apache 2.0 | 17–32B · high-memory workstation | 🇨🇳 China |
| Qwen3-VL-30B-A3B-InstructQwen / Alibaba | 30B-class sparse vision-language model | Apache 2.0 | 17–32B · high-memory workstation | 🇨🇳 China |
| Mistral Small 4 119B A6BMistral AI | 119B / 6.5B active · multimodal MoE | Apache 2.0 | Model-specific · large / specialized | 🇫🇷 France |
| Command A+ 05-2026Cohere / Cohere Labs | 218B / 25B active · multimodal MoE | Apache 2.0 | Model-specific · large / specialized | 🇨🇦 Canada |
| gpt-oss-20bOpenAI | 20B · compact reasoning | Apache 2.0 | 17–32B · high-memory workstation | 🇺🇸 United States |
| Gemma 3 27B ITGoogle DeepMind | 27B · multimodal | Gemma Terms | 17–32B · high-memory workstation | 🇺🇸 United States |
| Granite 4.2 8BIBM | 8B · 128K enterprise reasoning | Apache 2.0 | ≤8B · consumer/local | 🇺🇸 United States |
| Mistral Nemo Instruct 2407Mistral AI / NVIDIA | 12B-class · multilingual | Apache 2.0 | 9–16B · workstation/local | 🇫🇷 France |
Decision criteria
Benchmark realistic document lengths instead of only the published maximum.
Track KV-cache memory and prefill latency as context grows.
Test information placed at the beginning, middle and end of long prompts.
For RAG, compare retrieval-assisted prompts with full-context stuffing.
Deployment reality
Validate the exact checkpoint, precision or quantization, runtime, context length and concurrency target. Weight memory alone does not capture KV cache, runtime workspaces, multimodal encoders or distributed-serving overhead.
OWM keeps license, provider origin and data residency separate. A provider-country label is provenance metadata; the deployer determines where inference and connected services run.