Home›Compare›Architecture decision
Architecture decision · source-first decision page

Dense vs MoE Open-Weight Models

Compare dense and mixture-of-experts open-weight models by total parameters, active parameters, checkpoint memory, compute, serving complexity and workload fit.

Updated 1 Oct 20266 referenced modelsNo universal ranking
Direct answer

Dense and MoE parameter counts are not directly comparable. Dense models use the full model for each token; MoE models route a token through a subset of experts. Active parameters can reduce per-token compute, but the total checkpoint still affects storage, memory placement and distributed serving. For deployment, record both total and active parameters and benchmark the exact runtime rather than assuming that a 30B/3B-active MoE behaves like a dense 3B model.

01Total params
02Active params
03Memory placement
04Measured serving

Shortlist snapshot

ModelArchitecture / sizeLicenseOWM hardware classProvider origin
Qwen3-32BQwen / Alibaba32B · denseApache 2.017–32B · high-memory workstationChina
Qwen3-30B-A3BQwen / Alibaba30B / 3B active · MoEApache 2.017–32B · high-memory workstationChina
gpt-oss-20bOpenAI20B · compact reasoningApache 2.017–32B · high-memory workstationUnited States
Qwen3-235B-A22BQwen / Alibaba235B / 22B active · MoEApache 2.0Model-specific · large / specializedChina
DeepSeek-R1DeepSeekreasoning model · MoEMITModel-specific · large / specializedChina
Mistral Small 4 119B A6BMistral AI119B / 6.5B active · multimodal MoEApache 2.0Model-specific · large / specializedEU provider
Important: a shortlist is not a ranking. Eliminate incompatible models first, then benchmark the survivors on the exact workload.

What should drive the decision?

The common mistake is to treat active parameters as if they were the model’s full memory footprint. Expert weights that are inactive for a particular token still exist and must be stored somewhere accessible to the serving system.

MoE can improve compute efficiency at a given total capacity, but expert routing can introduce communication and load-balancing complexity. This becomes especially important when experts are split across devices.

Dense models are often simpler to reason about for memory estimation and runtime support. MoE models can be attractive when their quality/compute trade-off is favorable, but the operational stack must support the architecture well.

Models to evaluate

🇨🇳 Qwen / AlibabaApache 2.0

Qwen3-30B-A3B

Efficient reasoning and multilingual inference

30B / 3B active · MoE17–32B · high-memory workstation

30B total / about 3B active MoE reference point.

🇺🇸 OpenAIApache 2.0

gpt-oss-20b

Reasoning, local and edge-class inference

20B · compact reasoning17–32B · high-memory workstation

About 21B total / 3.6B active sparse reasoning model.

Hardware and runtime reality

Weight-only estimates are a starting point. Add KV cache, runtime workspaces, multimodal components, batching and concurrency headroom. For local inference, validate the exact quantized artifact. For server inference, measure time to first token, throughput and peak memory at target concurrency.

Long context can make an otherwise comfortable model exceed the practical memory budget. Test the longest realistic prompt and generation, not only a short loading test.

License, provider origin and Europe

Review the exact checkpoint license and any separate usage terms. Provider origin is supply-chain metadata, not an inference-location claim. If EU/EEA residency matters, map inference, RAG, embeddings, logs, telemetry, backups and subprocessors.

Evaluation checklist

Task qualityRepresentative prompts and hard cases.
ReliabilityTool errors, malformed output and regressions.
Latency + throughputMeasure target concurrency.
Peak memoryUse realistic context lengths.
License fitExact checkpoint and distribution model.
Data pathInference, retrieval, logs and backups.

Primary sources and related references