Detailed OWM Model Reference · verified 2026-09-27

Llama 4 Scout

Meta · Native multimodal Mixture-of-Experts · Llama 4 Community License

Direct answer: Llama 4 Scout is a useful case study in the difference between technical openness and licensing simplicity. Its weights are available, its 10M context is exceptional, and it is natively multimodal — but the checkpoint sits under Meta’s custom Llama 4 Community License rather than a standard permissive license. OWM therefore views it as technically portable but legally more conditional than Apache/MIT alternatives.

109B total · 17B active 10M context Llama 4 Community License Image + text → text/code
OWM View

Llama 4 Scout: what matters beyond the model card

Meta documents 109B total parameters with 17B activated, image and text input and a 10M context window. That context scale creates genuinely different system-design possibilities, but it also magnifies memory and serving complexity. The Llama ecosystem is broad, yet sovereignty analysis has to include the license obligations and not just the availability of weights.

Open Weight Models editorial view

Llama 4 Scout is a useful case study in the difference between technical openness and licensing simplicity. Its weights are available, its 10M context is exceptional, and it is natively multimodal — but the checkpoint sits under Meta’s custom Llama 4 Community License rather than a standard permissive license. OWM therefore views it as technically portable but legally more conditional than Apache/MIT alternatives.

Model facts

DeveloperMeta
Exact model IDmeta-llama/Llama-4-Scout-17B-16E-Instruct
Parameters109B total · 17B active
Context10M
ArchitectureNative multimodal Mixture-of-Experts
ModalitiesImage + text → text/code
LicenseLlama 4 Community License
AccessLicense acceptance / gated distribution

Runtime paths recorded by OWM: Meta reference implementation · Transformers. Runtime support is version-sensitive; a named runtime should not be read as a guarantee that every quantization, context size or feature works identically.

Why this model matters

Meta documents 109B total parameters with 17B activated, image and text input and a 10M context window. That context scale creates genuinely different system-design possibilities, but it also magnifies memory and serving complexity. The Llama ecosystem is broad, yet sovereignty analysis has to include the license obligations and not just the availability of weights.

OWM evaluates a model as infrastructure, not only as a benchmark entry. That means the exact checkpoint, license, runtime ecosystem, memory footprint, ability to move between providers and the quality of the evidence all matter alongside model capability.

Hardware reality

Meta states that the Llama 4 series requires at least four GPUs for full BF16 inference with the reference implementation. Quantized or alternative runtime paths may change that, but OWM does not substitute community conversions for the publisher’s documented full-precision guidance.

OWM deliberately separates raw-weight arithmetic, publisher guidance and measured runtime evidence. A model that theoretically fits into a memory budget can still fail in practice because of KV cache, runtime buffers, vision components, tensor-parallel overhead or concurrent requests.

License reality

The Llama 4 Community License grants broad rights subject to conditions. Redistribution includes notice and “Built with Llama” obligations, and additional terms apply for very large services above the license’s stated monthly-active-user threshold. That makes the license a first-class deployment fact, not a footnote.

This is an informational deployment summary, not legal advice. Production users should review the exact current license, usage policy, derivative-model terms and applicable law before shipping a product.

OWM Sovereignty Lens

OWM does not assign a single sovereignty score. We describe the layers separately because a model can be highly portable technically while remaining conditional legally — or permissively licensed while requiring infrastructure that limits practical choice.

Weight control
Strong but gated
The checkpoint can be obtained and retained after accepting the license.
License freedom
Conditional
The custom community license includes redistribution and large-user conditions.
Deployment control
Strong
Once obtained, the model can be self-hosted.
Runtime portability
Moderate to strong
The ecosystem is broad, although exact Llama 4 support varies by runtime.
Exit capability
Strong at provider layer
Infrastructure providers can be changed while retaining the checkpoint, subject to license compliance.

Runtime evidence

Evidence status

OWM currently records publisher and source-verified runtime information for Scout. Independent configuration-specific performance data should be added without converting it into an OWM-tested claim.

“Third-party measured” means the result was measured outside Open Weight Models and is shown with provenance. “OWM runtime tested” is reserved for configurations that OWM physically reproduces with an exact checkpoint, runtime version, hardware configuration, workload and date.

Change history

Initial snapshot

First OWM verification snapshot created. From this date forward, material changes to the model card, license, checkpoints, runtime support and deployment facts can be appended without reconstructing unobserved history.

Open the global OWM Change History →

Where Llama 4 Scout fits — and where it does not

Where it fits

  • Long-context research and applications where 10M-token capability is genuinely useful.
  • Multimodal systems that benefit from Meta’s Llama ecosystem.
  • Organizations comfortable reviewing and complying with the Llama 4 Community License.
  • Infrastructure teams with multi-GPU capacity.

Where it does not fit

  • Teams requiring a simple Apache/MIT licensing posture.
  • Single-consumer-GPU full-precision deployment.
  • Applications treating maximum context as free: extreme context lengths can dominate memory and latency.

Open-weight significance

The strategic value of this model is not simply that its weights can be downloaded. The important question is what those weights let an operator control: infrastructure, data location, runtime, adaptation and the ability to exit a provider relationship without discarding the model layer. Those freedoms remain bounded by the model’s license and the practical hardware required to run it.

Frequently asked questions

How large is Llama 4 Scout?

Meta documents 109B total parameters with 17B activated parameters.

What is the context window of Llama 4 Scout?

Meta lists a 10M-token context window for Scout.

Is Llama 4 Scout Apache 2.0?

No. It uses Meta’s Llama 4 Community License.

Can Llama 4 Scout be self-hosted?

Yes, subject to the license and substantial hardware requirements. Meta states at least four GPUs for full BF16 reference inference.

Primary sources and OWM data

Last verified by Open Weight Models: 2026-09-27. Facts can change as model repositories, licenses and runtime support evolve.

Related model references