Gold-standard Model Passport · full technical review

Llama 4 Maverick

A 400B-parameter natively multimodal Mixture-of-Experts model with 17B active parameters, 128 routed experts and a documented 1M-token base context window.

This passport separates the base model from its legal conditions and from each deployment implementation. That distinction matters for Maverick: weight access is broad but not unrestricted, the multimodal Llama 4 policy contains an EU developer restriction, and real context or feature limits differ substantially across self-hosted and managed deployments.

VERIFIEDOPEN WEIGHTSMULTIMODALMIXTURE-OF-EXPERTSCUSTOM LICENSECOMMERCIAL: CONDITIONAL
Source-firstPrimary Meta model card, license and policy separated from provider documentation.
License-awareOpen weight does not mean unrestricted use or permissive open-source licensing.
Deployment-specificContext, tools, regions and hardware are recorded per implementation.
Evidence datedFull technical and legal-source review: 29 September 2026.
Definition

What is Llama 4 Maverick?

Llama 4 Maverick is a natively multimodal open-weight AI model developed by Meta as part of the Llama 4 model family. It is designed to process both text and images and to generate multilingual text and code.

Maverick uses a Mixture-of-Experts (MoE) architecture with approximately 400 billion total parameters, while about 17 billion parameters are active for each token. The model contains 128 routed experts plus a shared expert and supports a documented context window of up to 1 million tokens.

Unlike a conventional dense 17B model, Llama 4 Maverick still requires infrastructure capable of storing and serving its much larger total parameter set. The sparse MoE architecture reduces the amount of computation used for each token, but it does not reduce the complete model to the hardware footprint of a 17B model.

Meta released Llama 4 Maverick on 5 April 2025 as one of the first Llama models combining native multimodality with a Mixture-of-Experts architecture. Official weights are available for independent deployment under the Llama 4 Community License Agreement.

In practical terms, Llama 4 Maverick is intended for workloads such as multimodal assistants, image understanding, document analysis, coding, multilingual generation, long-context processing and AI systems that require self-hosted or managed deployment options.

Meta Llama 4 model card ↗
Executive summary

Why Maverick is technically unusual.

Llama 4 Maverick combines the storage footprint of a very large model with the per-token compute characteristics of a sparse Mixture-of-Experts system. It is therefore misleading to treat “17B active parameters” as if Maverick were a conventional 17B model.

Meta documents approximately 400 billion total parameters, while only about 17 billion parameters are activated for an individual token. The MoE layers contain 128 routed experts plus a shared expert; each token passes through the shared expert and one routed expert. The sparse route lowers active computation, but the whole parameter set still has to be available to the serving system. That is why deployment remains datacenter-class even though the active count looks comparatively small.

Maverick is also natively multimodal. Meta uses early fusion so image and text tokens are processed in a unified model backbone rather than attaching vision only as a late external component. The official model card lists multilingual text and image input, multilingual text and code output, a 1M context window and an August 2024 knowledge cutoff.

Legally, Maverick is not an Apache-2.0- or MIT-style release. It is distributed under the Llama 4 Community License Agreement, with an incorporated Acceptable Use Policy. The license grants broad rights to use, modify and redistribute subject to conditions, while the policy adds prohibited-use rules and, for Llama 4 multimodal models, an important EU restriction for certain developers. This passport therefore marks commercial use as conditional, not simply “yes”.

Meta model card ↗Meta architecture announcement ↗
OpenWeightModels classification: Maverick is an open-weight model because official weights are available for independent deployment. That classification says nothing by itself about whether its license is OSI-approved, unrestricted, region-neutral or permissive.
Model Passport

Core facts at a glance.

These values describe the base Maverick model unless a deployment provider is explicitly named. Provider-specific limits appear later and should not be silently inherited from the base model.

DeveloperMetaLlama 4
Model typeAutoregressive MoEnative multimodality
Total parameters~400Bstored parameter set
Active parameters17Bper-token active scale
Routed experts128plus shared expert
Base context1M tokensMeta model card
Training tokens~22Tmultimodal pretraining
Knowledge cutoffAug 2024static model
Supported languages12explicitly listed
Pretraining languages200broader training coverage
Official precisionsBF16 + FP8quantized release available
Commercial useConditionalcustom license + policy
Architecture

400B stored. 17B active.

Meta describes Llama 4 as an autoregressive architecture using alternating dense and Mixture-of-Experts layers. In Maverick's MoE layers, routing sends each token through a shared expert and one of 128 routed experts. The design aims to preserve the representational capacity of a very large parameter pool without activating that entire pool for every token.

This distinction affects both performance and infrastructure. Active parameters are relevant to compute performed per token, but total parameters remain relevant to weight storage, model loading, sharding and memory topology. A capacity plan based solely on 17B would therefore substantially understate Maverick's deployment footprint.

InputText + image tokens
→
Unified backboneEarly-fusion multimodal transformer
→
MoE routingShared expert + 1 of 128 routed experts
01

Mixture-of-Experts

Only a subset of parameters participates in each token's forward pass. This reduces active compute relative to a dense model with the same total parameter count, but it does not eliminate the need to store and serve the complete expert set.

02

Early fusion

Text and visual tokens are integrated into the same model backbone. Meta says its vision encoder is based on MetaCLIP and was adapted in conjunction with a frozen Llama model before multimodal joint training.

03

Long context

Meta documents 1M tokens for Maverick. Long context has direct implications for KV-cache memory, latency and serving economics, so the model maximum should never be treated as a cost-free default operating point.

Practical reading of “17B active”: it is primarily a compute-routing fact, not a statement that the whole model can be hosted like a dense 17B checkpoint. Meta itself says all parameters remain stored in memory while only a subset is activated during serving.
Meta Llama 4 architecture ↗
Training & post-training

Multimodal from pretraining onward.

Meta reports roughly 22 trillion multimodal pretraining tokens for Maverick from a mixture of publicly available data, licensed data and information from Meta products and services, including publicly shared Instagram and Facebook posts and interactions with Meta AI. The model card lists August 2024 as the data freshness cutoff.

The broader Llama 4 pretraining effort covered 200 languages, with more than 100 languages receiving over one billion tokens each. Meta explicitly lists 12 languages as supported for Maverick deployments, so broader pretraining coverage should not be interpreted as equivalent support quality.

For post-training, Meta describes a pipeline of lightweight supervised fine-tuning, online reinforcement learning and lightweight direct preference optimization. It also says Maverick was co-distilled from the larger Llama 4 Behemoth teacher during pretraining.

Pretraining

~22T multimodal tokens; text, image and video-related data; broad multilingual corpus. Meta says Llama 4 models were pretrained with up to 48 images in multimodal sequences.

Post-training

Lightweight SFT → online RL → lightweight DPO, with a curriculum intended to balance multimodal, reasoning and conversational capability.

Distillation

Meta reports co-distillation from Llama 4 Behemoth, a much larger teacher model, to improve end-task quality while keeping Maverick's active computation lower.

Training compute disclosure: Meta reports 2.38 million H100-80GB GPU-hours for Maverick pretraining and 645 tons of location-based CO₂e for that training run. This is a developer-reported environmental disclosure, not an independent audit.
Official model card ↗Meta release blog ↗
Capabilities

What Maverick is designed to do.

Capabilities below separate base-model intent from provider-added features. Tool calling, structured outputs and guardrails can change depending on the serving platform even though the underlying weights remain the same.

Text & instruction following

The Instruct checkpoint is intended for assistant-like chat, generation, summarization, multilingual workflows and general instruction following.

Image understanding

Meta lists visual recognition, image reasoning, captioning and visual question answering. The model is text-output only; it does not generate images.

Multi-image reasoning

Meta's model card says the family was tested for image understanding with up to five input images and recommends additional application-specific testing beyond documented conditions.

Coding

The model card includes code benchmarks and describes multilingual text and code output. Coding capability is part of the core model rather than a cloud-only extension.

Long-context work

The 1M base context supports large-document and code-context scenarios, but provider implementations may expose lower practical limits and long contexts increase memory and latency costs.

Tool use

Do not treat tool calling as a universal Maverick property. AWS documents client-side tool calling and Agents; Google documents function calling; current Microsoft Foundry documentation lists tool calling as unavailable for its Maverick offering.

Evaluation evidence

Selected Meta-reported benchmarks.

The table below is intentionally limited to a small set of measurements from Meta's official model card. Meta states that its reported evaluations were conducted on BF16 models even though quantized checkpoints are also available. Scores should not be treated as universal real-world performance guarantees.

MMLU-Pro · Instruct80.50-shot macro average accuracy
GPQA Diamond69.80-shot accuracy
LiveCodeBench43.40-shot pass@1, stated test window
MMMU73.40-shot image reasoning accuracy
AreaBenchmarkMetricMaverickInterpretation
Reasoning & knowledgeMMLU Promacro_avg / accuracy80.5Instruction-tuned, 0-shot in Meta's evaluation.
Science reasoningGPQA Diamondaccuracy69.8High-difficulty factual/reasoning benchmark.
CodingLiveCodeBenchpass@143.4Meta reports a dated evaluation window; benchmark versions matter.
Image reasoningMMMUaccuracy73.4Measures multimodal reasoning across academic domains.
Image understandingChartQArelaxed accuracy90.0Chart-oriented visual question answering.
Multilingual mathMGSMaverage / exact match92.3Multilingual grade-school math evaluation.
Benchmark policy: OpenWeightModels records the evaluator, checkpoint/precision where known, metric and test conditions. It does not convert heterogeneous benchmark results into a single overall model score.
Meta model-card benchmark table ↗
License intelligence

Broad rights, material conditions.

Maverick uses the Llama 4 Community License Agreement, not a standard permissive software license. Meta grants a limited, worldwide, royalty-free license to use, reproduce, distribute, create derivative works and modify the Llama Materials, but redistribution, attribution, scale and policy conditions apply.

The Acceptable Use Policy is explicitly incorporated into the agreement. OpenWeightModels therefore treats the license document and the policy document as two separate but jointly relevant sources.

Official license

Llama 4 Community License Agreement

Custom commercial/community license · effective 5 April 2025 · accompanied by incorporated Acceptable Use Policy.

QuestionClassificationWhat the source says in practice
Access / use weightsAllowed*License grants rights to use Llama Materials, subject to the agreement and incorporated policy.
Commercial useConditionalCommercial/research use is intended, but scale, policy and regional conditions can materially affect eligibility.
Modification / fine-tuningAllowed*Modification and derivative works are included in the grant, subject to the terms.
RedistributionConditionalAgreement copy, “Built with Llama” display and Notice-file attribution requirements apply when distributing/making materials or containing products/services available.
Derived AI model namingConditionalIf Llama materials or outputs are used to create/train/fine-tune/improve another AI model that is distributed or made available, the agreement requires “Llama” at the beginning of that model name.
700M MAU clauseSpecial termIf the licensee or affiliates exceeded 700M monthly active users in the month preceding the Llama 4 release, a separate Meta license is required before exercising rights.
Acceptable Use PolicyMandatoryThe policy is incorporated by reference and contains prohibited-use categories plus the multimodal EU developer restriction.
OSI-approved open sourceNo standard OSS licenseOpenWeightModels classifies the weights as open-weight but does not classify the license as a standard permissive OSI software license.
Redistribution: Meta requires a copy of the agreement with distributed Llama Materials, prominent “Built with Llama” attribution on a related surface, and a specified Notice-file attribution for copies of Llama Materials.
Large-platform threshold: The 700M-MAU provision is tied to the Llama 4 release date and preceding calendar month. It is a special eligibility condition, not a recurring generic “over 700M users” fee tier.
EU multimodal restriction: Meta's current Llama 4 Acceptable Use Policy states that the Section 1(a) rights for multimodal Llama 4 models are not granted to an individual domiciled in, or a company with its principal place of business in, the European Union. The policy separately says this restriction does not apply to end users of a product or service that incorporates such multimodal models. This means developer/license eligibility and end-user access are not the same question.
OpenWeightModels interpretation: “Commercial use: conditional” is more accurate than a binary “yes.” A typical developer outside the EU may have broad commercial rights under the agreement, but redistribution obligations, the AUP, the 700M-MAU clause and geography can change the answer for a specific organization or deployment.
Official Llama 4 license ↗Official Llama 4 Acceptable Use Policy ↗
Deployment intelligence

One model, multiple operating envelopes.

The phrase “supports 1M context” is incomplete without naming the implementation. Maverick's base model is documented at 1M, NVIDIA NIM supports 1M on H200 but 430K on H100, Google currently documents 524,288 on its managed endpoint, while AWS documents 1M. Provider-specific features such as tool calling also differ.

OpenWeightModels therefore stores deployment facts as child records of the model rather than flattening them into a single global feature list.

RouteControlContext / outputModalitiesNotable feature
Direct BF16 / FP8 weightsHighestBase model: 1M documentedText + image → textFull operator responsibility for serving, memory and policy compliance.
Hugging Face + vLLM/SGLangHighImplementation-dependentText + image → textOfficial model page documents Transformers, vLLM and SGLang launch paths.
NVIDIA NIMHigh / packagedH200: 1M · H100: 430KVision-language servingOfficial FP8 checkpoint only; validated 8-GPU node configurations.
Amazon BedrockManaged1M context · 8K max outputText + image → textStreaming, Guardrails, client-side tool calling, Flows and Agents documented.
Google managed endpointManaged524,288 context · 8,192 outputText/code/image → textFunction calling and structured output documented; model availability in us-east5.
Microsoft FoundryManaged1M documented input contextText + images → textGlobal Standard deployment; current Microsoft listing says tool calling: No.

Direct weights

self-host

Meta releases both BF16 and official FP8 Maverick checkpoints. Direct deployment offers the most control but shifts model serving, scaling, observability, safeguards and legal compliance to the operator.

ContextUp to 1M base documentationPrecisionBF16 / FP8 officialDifficultyVery high
Meta model card ↗

Hugging Face runtimes

self-host

The official model page documents Transformers loading plus vLLM and SGLang serving, including OpenAI-compatible API examples for the serving frameworks.

RuntimesTransformers · vLLM · SGLangAPI shapeOpenAI-compatible in serving examplesControlHigh
Official HF Instruct model ↗

NVIDIA NIM

packaged

NVIDIA validates Maverick only on homogeneous eight-GPU H100/H200 nodes in its current NIM support matrix. Generic configurations are not supported for this profile.

H200 context1,000,000H100 context430,000CheckpointMeta official FP8 only
NVIDIA support matrix ↗

Amazon Bedrock

managed

AWS lists Maverick as an active Bedrock model with text/image input and text output. AWS also applies geofencing to Llama 4 access based on account country and request source.

Context1MMax output8KFeaturesStreaming · Guardrails · tools · Agents
AWS model card ↗AWS geofencing note ↗

Google Cloud

managed

Google's current managed Maverick record exposes a materially smaller context envelope than Meta's base-model maximum and adds platform features such as function calling and structured output.

Model IDllama-4-maverick-17b-128e-instruct-maasContext524,288Regionus-east5
Google model record ↗

Microsoft Foundry

managed

Microsoft currently lists the FP8 Maverick Instruct model as a Foundry model sold directly by Azure with Global Standard deployment. Its current capability table does not advertise tool calling for this model.

InputText + imagesContext1M documentedTool callingNo (current listing)
Microsoft Foundry model list ↗
Why provider records matter: the model identity is the same, but the operating contract is not. A cloud endpoint can expose a smaller context window, different tool support, different safety features, different regional availability and a separate lifecycle from the underlying downloadable weights.
Hardware & quantization

Active parameters do not equal memory footprint.

Meta releases Maverick in BF16 and FP8. It states that the FP8 weights can fit on a single H100 DGX host while maintaining quality. NVIDIA's production NIM profile is more explicit: it supports a node of eight H100 SXM, H100 NVL, H200 SXM or H200 NVL GPUs and uses Meta's official FP8 checkpoint only.

Long context further changes memory requirements because KV-cache usage grows with sequence length, batch size and implementation. For this reason, OpenWeightModels does not publish one universal “minimum VRAM” number for Maverick.

BF16

Official Meta weight format and the precision Meta says it used for its reported evaluations. It provides a high-quality reference point but carries the largest raw weight-memory requirement of the official releases.

FP8

Official quantized release. Meta says the FP8 weights fit on one H100 DGX host. NVIDIA NIM uses this official FP8 checkpoint for its validated Maverick serving profile.

On-the-fly INT4

Meta also provides code for on-the-fly INT4 quantization. That is distinct from saying every INT4 community artifact has identical behavior, provenance or quality.

Hardware interpretation: “single H100 host” is not “single H100 GPU.” A DGX H100 is an eight-GPU system. This distinction should remain explicit in technical summaries because collapsing “host” into “GPU” radically understates infrastructure needs.
Meta quantization notes ↗NVIDIA NIM hardware matrix ↗
Weights & variants

Know which Maverick you are actually deploying.

Model-family names are not deployment identifiers. Precision and tuning state matter for reproducibility, memory planning and comparison.

VariantPurposePrecisionStatusNotes
Llama-4-Maverick-17B-128E-OriginalOriginal/pretrained-style releaseBF16OfficialUseful when adapting the base model rather than consuming the instruction-tuned behavior.
Llama-4-Maverick-17B-128E-InstructAssistant / visual reasoningBF16OfficialPrimary instruction-tuned checkpoint documented for chat and visual reasoning.
Llama-4-Maverick-17B-128E-Instruct-FP8Reduced-memory servingFP8OfficialOfficial quantized release; used by NVIDIA NIM's validated Maverick profile.
Community GGUF / other quantizationsAlternative runtimesVariesCommunityShould be tracked separately by publisher, quantization recipe and source; not equivalent to Meta's official checkpoint.
Limitations & selection notes

Where the headline numbers can mislead.

A useful model passport should make deployment risks and interpretive limits as visible as capabilities.

It is not a “17B hardware model”

The 17B figure is the active parameter count. Maverick still contains roughly 400B total parameters, making model storage and orchestration a large-system problem.

1M context is not universal

The base model is documented at 1M, but Google exposes 524,288 on its managed endpoint and NVIDIA limits H100 NIM deployments to 430K. Provider context must be recorded separately.

Tool support is provider-specific

AWS and Google expose tool/function-oriented capabilities around their Maverick services; current Microsoft Foundry documentation lists tool calling as unavailable for its Maverick model.

License is not region-neutral

The incorporated AUP contains a specific EU developer restriction for Llama 4 multimodal models while carving out end users of incorporated products/services. This deserves explicit legal review for EU organizations.

Static knowledge

Meta lists August 2024 as the knowledge cutoff. Retrieval or other external-data architecture may be needed for newer factual information.

Benchmarks are not deployment guarantees

Meta says its reported evaluations were run on BF16 models. Provider quantization, serving parameters, prompts, safety layers and context lengths can change application-level behavior.

Legal note: OpenWeightModels summarizes published terms for technical comparison and provenance. It is not legal advice. Organizations should review the current official agreement, use policy and provider terms for their jurisdiction and intended use.
Sources & verification

Evidence behind the passport.

Last full review: 29 September 2026. Architecture, training, benchmark and weight claims are grounded primarily in Meta sources. Deployment-specific claims are grounded in the documentation of the named platform rather than inferred from the base model.

Change history

Passport revisions.

Material source changes should update both the affected field and this history instead of silently overwriting the record.

Gold-standard passport created. Expanded architecture, training, capabilities, benchmarks, detailed license analysis, provider-level deployment records, hardware/quantization, variants and limitations. EU restriction source corrected to the incorporated Acceptable Use Policy rather than the core license text.

Initial Model Passport created with license and deployment overview.

FAQ

Common interpretation questions.

Is Llama 4 Maverick open source?

OpenWeightModels calls it open-weight because official weights are available. Its custom Llama 4 Community License and incorporated use policy differ from standard permissive software licenses, so “open weight” is the more precise classification for this registry.

Can Llama 4 Maverick be used commercially?

Meta's model card intends Llama 4 for commercial and research use, but the answer is conditional on the Community License and AUP. Important conditions include redistribution obligations, the 700M-MAU clause and the multimodal EU developer restriction.

Can Llama 4 Maverick run on a single H100 GPU?

Do not confuse a single H100 GPU with an H100 DGX host. Meta says the FP8 checkpoint fits on a single H100 DGX host; NVIDIA's validated NIM profile uses eight H100 or H200 GPUs.

What is the context window of Llama 4 Maverick?

Meta documents a base context window of up to 1 million tokens for Llama 4 Maverick. Provider implementations can expose smaller limits: Google's current managed Maverick endpoint documents 524,288 tokens, while NVIDIA NIM documents 430,000 on H100 and 1 million on H200.

Does Maverick support tool calling?

There is no single provider-independent yes/no answer in practice. AWS documents client-side tool calling, Google documents function calling, while current Microsoft Foundry documentation lists tool calling as unavailable for its Maverick offering.