Detailed OWM Model Reference · verified 2026-09-27

Mistral Small 4

Mistral AI · MoE · 128 experts · 4 active · Apache 2.0

Direct answer: Mistral Small 4 is an unusually dense package of deployment features: multimodal input, reasoning and non-reasoning modes, function calling, a 256K context window, MoE efficiency and Apache 2.0 licensing. OWM sees the official NVFP4 checkpoint as especially important because it turns quantization from a community afterthought into a publisher-supported deployment path.

119B total · 6.5B active/token 256K context Apache 2.0 Image + text → text
OWM View

Mistral Small 4: what matters beyond the model card

The model uses 119B total parameters but activates 6.5B per token. Mistral documents 128 experts with four active, image and text input, reasoning controls and function calling. The model’s value for infrastructure teams is therefore not one benchmark number but the combination of a large sparse architecture with an explicit low-precision deployment option.

Open Weight Models editorial view

Mistral Small 4 is an unusually dense package of deployment features: multimodal input, reasoning and non-reasoning modes, function calling, a 256K context window, MoE efficiency and Apache 2.0 licensing. OWM sees the official NVFP4 checkpoint as especially important because it turns quantization from a community afterthought into a publisher-supported deployment path.

Model facts

DeveloperMistral AI
Exact model IDmistralai/Mistral-Small-4-119B-2603
Parameters119B total · 6.5B active/token
Context256K
ArchitectureMoE · 128 experts · 4 active
ModalitiesImage + text → text
LicenseApache 2.0
AccessDirect open-weight download

Runtime paths recorded by OWM: Transformers/Mistral4 · vLLM. Runtime support is version-sensitive; a named runtime should not be read as a guarantee that every quantization, context size or feature works identically.

Why this model matters

The model uses 119B total parameters but activates 6.5B per token. Mistral documents 128 experts with four active, image and text input, reasoning controls and function calling. The model’s value for infrastructure teams is therefore not one benchmark number but the combination of a large sparse architecture with an explicit low-precision deployment option.

OWM evaluates a model as infrastructure, not only as a benchmark entry. That means the exact checkpoint, license, runtime ecosystem, memory footprint, ability to move between providers and the quality of the evidence all matter alongside model capability.

Hardware reality

The main repository is hundreds of gigabytes at high precision. A simplistic raw 4-bit calculation for 119B parameters is still about 59.5 GB before overhead. Mistral’s official NVFP4 checkpoint is therefore strategically relevant: it offers a vendor-supported low-precision route rather than relying only on third-party conversions.

OWM deliberately separates raw-weight arithmetic, publisher guidance and measured runtime evidence. A model that theoretically fits into a memory budget can still fail in practice because of KV cache, runtime buffers, vision components, tensor-parallel overhead or concurrent requests.

License reality

Mistral publishes Mistral Small 4 under Apache 2.0 for commercial and non-commercial use. OWM treats that standard permissive license as a major advantage for organizations that want to build internal or customer-facing systems without a custom model license.

This is an informational deployment summary, not legal advice. Production users should review the exact current license, usage policy, derivative-model terms and applicable law before shipping a product.

OWM Sovereignty Lens

OWM does not assign a single sovereignty score. We describe the layers separately because a model can be highly portable technically while remaining conditional legally — or permissively licensed while requiring infrastructure that limits practical choice.

Weight control
Strong
The model and an official low-precision checkpoint are downloadable.
License freedom
Strong
Apache 2.0 is broadly permissive.
Deployment control
Strong but hardware-sensitive
Self-hosting is possible, though the full checkpoint remains large.
Runtime portability
Moderate to strong
Current vLLM and Transformers support is documented; ecosystem breadth is still narrower than older architectures.
Exit capability
Strong
The Apache-licensed checkpoint can move with the operator.

Runtime evidence

Evidence status

The current Passport records source-verified support for vLLM and the Mistral4 Transformers integration. OWM has not yet reproduced a standardized runtime benchmark for this exact checkpoint.

“Third-party measured” means the result was measured outside Open Weight Models and is shown with provenance. “OWM runtime tested” is reserved for configurations that OWM physically reproduces with an exact checkpoint, runtime version, hardware configuration, workload and date.

Change history

Initial snapshot

First OWM verification snapshot created. From this date forward, material changes to the model card, license, checkpoints, runtime support and deployment facts can be appended without reconstructing unobserved history.

Open the global OWM Change History →

Where Mistral Small 4 fits — and where it does not

Where it fits

  • Enterprise agent systems needing vision, function calling and long context.
  • Teams with GPU infrastructure that can benefit from an official NVFP4 checkpoint.
  • Organizations preferring Apache 2.0 for commercial deployment.
  • Serving stacks built around current vLLM or Transformers/Mistral4 support.

Where it does not fit

  • Small consumer systems expecting a 6.5B-active model to have 6.5B storage requirements.
  • Very long contexts without capacity planning for KV cache.
  • Teams on older inference stacks that do not yet support the Mistral4 architecture.

Open-weight significance

The strategic value of this model is not simply that its weights can be downloaded. The important question is what those weights let an operator control: infrastructure, data location, runtime, adaptation and the ability to exit a provider relationship without discarding the model layer. Those freedoms remain bounded by the model’s license and the practical hardware required to run it.

Frequently asked questions

How many parameters does Mistral Small 4 have?

Mistral documents 119B total parameters with 6.5B activated per token.

What is the context length of Mistral Small 4?

The model card documents a 256K context window.

Is Mistral Small 4 multimodal?

Yes. It accepts text and image input and produces text output.

Does Mistral Small 4 have an official 4-bit checkpoint?

Mistral publishes an NVFP4 checkpoint, providing an official lower-precision deployment path.

Primary sources and OWM data

Last verified by Open Weight Models: 2026-09-27. Facts can change as model repositories, licenses and runtime support evolve.

Related model references