Shortlist snapshot
| Model | Architecture / size | License | OWM hardware class | Provider origin |
|---|---|---|---|---|
| Qwen3-32BQwen / Alibaba | 32B · dense | Apache 2.0 | 17–32B · high-memory workstation | China |
| Qwen3-30B-A3BQwen / Alibaba | 30B / 3B active · MoE | Apache 2.0 | 17–32B · high-memory workstation | China |
| gpt-oss-20bOpenAI | 20B · compact reasoning | Apache 2.0 | 17–32B · high-memory workstation | United States |
| Qwen3-235B-A22BQwen / Alibaba | 235B / 22B active · MoE | Apache 2.0 | Model-specific · large / specialized | China |
| DeepSeek-R1DeepSeek | reasoning model · MoE | MIT | Model-specific · large / specialized | China |
| Mistral Small 4 119B A6BMistral AI | 119B / 6.5B active · multimodal MoE | Apache 2.0 | Model-specific · large / specialized | EU provider |
What should drive the decision?
The common mistake is to treat active parameters as if they were the model’s full memory footprint. Expert weights that are inactive for a particular token still exist and must be stored somewhere accessible to the serving system.
MoE can improve compute efficiency at a given total capacity, but expert routing can introduce communication and load-balancing complexity. This becomes especially important when experts are split across devices.
Dense models are often simpler to reason about for memory estimation and runtime support. MoE models can be attractive when their quality/compute trade-off is favorable, but the operational stack must support the architecture well.
Models to evaluate
Qwen3-32B
Reasoning, multilingual, general-purpose
Dense reference point in the 32B class.
Qwen3-30B-A3B
Efficient reasoning and multilingual inference
30B total / about 3B active MoE reference point.
gpt-oss-20b
Reasoning, local and edge-class inference
About 21B total / 3.6B active sparse reasoning model.
Qwen3-235B-A22B
Reasoning, multilingual, general-purpose
235B total / 22B active large MoE.
DeepSeek-R1
Reasoning, mathematics and coding
671B total / 37B active large MoE per publisher model card.
Mistral Small 4 119B A6B
Instruction, reasoning, coding, agents and vision
119B total / 6.5B active multimodal MoE in the OWM registry.
Hardware and runtime reality
Weight-only estimates are a starting point. Add KV cache, runtime workspaces, multimodal components, batching and concurrency headroom. For local inference, validate the exact quantized artifact. For server inference, measure time to first token, throughput and peak memory at target concurrency.
Long context can make an otherwise comfortable model exceed the practical memory budget. Test the longest realistic prompt and generation, not only a short loading test.
License, provider origin and Europe
Review the exact checkpoint license and any separate usage terms. Provider origin is supply-chain metadata, not an inference-location claim. If EU/EEA residency matters, map inference, RAG, embeddings, logs, telemetry, backups and subprocessors.