Open Weight Models
Home›Knowledge›Architecture
Architecture · source-first reference

AI Model Parameters Explained

Parameter count is one of the most visible model labels and one of the easiest to misuse. It helps estimate storage and memory, but it is not a universal quality score and it means something different for dense and sparse MoE models.

Direct answer

A model parameter is a learned numerical value used by the network during inference. A “7B” model has roughly seven billion parameters. In a dense model, most model parameters participate in each token’s forward pass. In a sparse Mixture-of-Experts model, the checkpoint can contain far more total parameters than the number active per token. Parameter count helps estimate model-weight memory, but quality, latency and total runtime memory depend on much more than the headline number.

BBillion parameters
Dense modelMost weights used per token
MoETotal ≠ active
Not a scoreBigger ≠ automatically better
Parameter scale and deployment intuition
<1Bedge / tiny
1–8Blocal-friendly
9–32Bworkstation
33–80Bhigh-memory
>80Bserver / MoE

These are orientation bands, not hardware guarantees. Precision, architecture, context and runtime can change the real requirement substantially.

What a parameter actually is

Neural networks learn numerical values that transform one representation into another. These values are stored in tensors and are adjusted during training. The final learned values are the model parameters or weights.

The parameter count therefore describes the size of the learned function in a very literal storage sense. If a model has 8 billion parameters and every parameter is stored using 16 bits, the raw weight tensor is on the order of 16 GB before considering metadata, runtime overhead, caches and temporary allocations.

This arithmetic is useful, but it is only the starting point for deployment planning.

Dense models: the intuitive case

In a dense transformer, the same network layers are generally executed for every token. A 7B dense model therefore has approximately seven billion learned parameters in its checkpoint, and the inference process uses the dense network for each token.

Dense parameter count is useful for comparing approximate checkpoint scale, but two equally sized models can have very different capabilities because training data, architecture, tokenizer, post-training, context design and optimization quality differ.

Do not convert parameter count into a quality score. It is a size/architecture fact, not a benchmark result.

MoE models: total parameters versus active parameters

Mixture-of-Experts models introduce sparse expert layers. A router selects a subset of experts for each token. This creates two important numbers:

  • Total parameters: all model weights stored in the checkpoint.
  • Active parameters per token: the approximate subset used for the computation of a token.

A 30B MoE with roughly 3B active parameters can have compute characteristics closer to a much smaller dense model for parts of the forward pass, but it still needs access to the expert weights. That is why “3B active” must not be used as if it were a 3B checkpoint size.

From parameter count to weight memory

Storage precisionApprox. bytes/parameter8B model weights32B model weights
FP324~32 GB~128 GB
FP16 / BF162~16 GB~64 GB
8-bit~1 + overhead~8 GB + overhead~32 GB + overhead
4-bit~0.5 + overhead~4 GB + overhead~16 GB + overhead

The table estimates weight storage only. Real inference memory also includes KV cache, runtime buffers, allocator overhead, multimodal encoders and sometimes multiple model components.

Why parameter count is a weak quality proxy

Parameter count became a convenient shorthand during the era when scaling larger dense transformers often correlated with capability improvements. But modern model design complicates that relationship. Distillation, synthetic-data post-training, sparse architectures, better tokenizers, reinforcement learning and specialized training can allow smaller models to outperform larger ones on particular tasks.

Use parameters to answer questions such as “How large is this checkpoint?” and “What hardware class should I investigate?” Use task-specific evaluation to answer “Is this model good enough for my workload?”

How to read model names without being misled

  1. Identify whether the model is dense or MoE.
  2. For MoE, record both total and active parameter counts.
  3. Identify the published precision or available quantizations.
  4. Check modality: vision encoders or audio components can add memory beyond the language backbone.
  5. Check context length and expected KV-cache requirements.
  6. Then map the checkpoint to a realistic runtime and hardware configuration.

Frequently asked questions

Does a 70B model always outperform a 7B model?

No. Model quality depends on architecture, training, post-training and the task. Parameter count is not a universal performance ranking.

How do I estimate raw weight memory?

As a rough first pass, multiply parameter count by bytes per parameter: FP32≈4, FP16/BF16≈2, 8-bit≈1 and 4-bit≈0.5, then add format and runtime overhead.

What does active parameters mean in a MoE model?

It is the subset of expert parameters selected for computation for a token. It does not mean the unused experts disappear from checkpoint storage.

Why can two 8B models require different memory?

Different architectures, vocabularies, attention designs, context lengths, multimodal encoders, quantization formats and runtimes can all change the real footprint.

Primary sources and technical references

OpenWeightModels prefers publisher documentation, standards bodies, official repositories and original research papers. The source material remains authoritative where it changes.

Continue the knowledge path