A Mixture-of-Experts (MoE) model contains multiple expert sub-networks and a router that selects a small subset of them for each token. This creates a sparse computation pattern: total parameters describe the full checkpoint, while active parameters describe the subset used per token. Sparse activation can reduce compute relative to a dense model of the same total size, but memory, communication and runtime support remain major deployment constraints.
Why MoE exists
Dense models use the same feed-forward parameters for every token. As models scale, that makes each token increasingly expensive. MoE takes another path: add more expert capacity but use a routing function so only selected experts process each token.
The Switch Transformer paper is a foundational example of this approach. It describes sparse activation as a way to increase model parameter count while keeping per-example computational cost more manageable than a similarly sized dense network.
Modern open-weight MoE releases use variations of this pattern and can contain hundreds of billions or even trillions of total parameters.
Total parameters and active parameters are different deployment facts
Suppose a model advertises 120B total parameters and 12B active parameters. The 12B number is useful for understanding how much expert computation is used per token, but the 120B number remains critical for checkpoint storage and model distribution.
Many deployment mistakes come from reading the active number as if the model “fits like a 12B dense model.” It usually does not. All expert weights must be available somewhere in the serving system, even if each token only touches a fraction of them.
| Number | What it describes | Useful for |
|---|---|---|
| Total parameters | Full learned parameter pool | Checkpoint scale, weight memory, distribution |
| Active parameters | Subset participating per token | Compute intuition, FLOPs, possible throughput |
| Experts / experts selected | Routing topology | Runtime support and communication behavior |
Why sparse compute does not remove the memory problem
An MoE server needs access to expert weights. On a single accelerator, that can exceed local VRAM. On multi-GPU systems, experts may be sharded across devices. That introduces communication: the router sends token representations to the relevant experts and returns their results.
Therefore, an MoE can be compute-efficient per token while still being memory-heavy and network-sensitive. High-bandwidth interconnects, expert-parallel support and good load balancing can matter as much as headline FLOPs.
Quantization can reduce weight memory, but runtime compatibility and expert-kernel quality become especially important for sparse models.
Routing, load balance and expert utilization
The router learns which experts to select. If routing becomes unbalanced, some experts can receive disproportionately many tokens, creating hot spots. Training methods often include mechanisms that encourage more balanced use of experts.
At inference time, routing patterns affect batching and hardware utilization. A runtime optimized for dense matrix multiplication may not automatically deliver good performance on a sparse expert model. This is why model publishers often recommend specific serving stacks or kernel versions for large MoE releases.
When MoE is attractive
MoE is attractive when a publisher wants high model capacity without activating every parameter for every token. For users, it can offer a strong capability-to-compute ratio — particularly in server environments that support the architecture well.
For local users, the decision is more nuanced. A sparse model with low active parameters can still have a very large checkpoint. If the entire model does not fit comfortably into RAM/VRAM, offloading and memory bandwidth can dominate performance.
Read the model card in this order
Total parameters → active parameters → experts selected → published precision → runtime support → hardware guidance. Skipping directly to “active parameters” produces misleading deployment expectations.
How MoE changes model comparison
Comparing a dense 32B model with a 120B/12B-active MoE is not a simple 32-versus-12 comparison. The dense model may be easier to fit and serve predictably; the MoE may offer greater parameter capacity with lower per-token expert compute but require more storage and more sophisticated serving.
The right comparison therefore includes quality on your workload, total memory, active compute, runtime maturity, throughput at your batch size, context length and operational complexity.
Frequently asked questions
Is a 30B model with 3B active parameters basically a 3B model?
No. The active count describes sparse per-token computation, while the full checkpoint still contains the expert weights and usually requires much more memory than a 3B dense model.
Why are MoE models faster than dense models of the same total size?
They can be because only a subset of experts is activated per token. Real speed also depends on routing, batching, memory bandwidth, kernels and multi-GPU communication.
Can MoE models be quantized?
Often yes, but support varies by architecture and runtime. Quantization reduces weight memory but does not eliminate expert-routing complexity.
What should I compare between MoE models?
Total parameters, active parameters, number of experts and selected experts, context, precision, runtime support, measured throughput and hardware topology.
Apply this knowledge
Move from the concept to a concrete deployment shortlist.
Primary sources and technical references
OpenWeightModels prefers publisher documentation, standards bodies, official repositories and original research papers. The source material remains authoritative where it changes.