Llama 4 Maverick is a natively multimodal open-weight AI model developed by Meta as part of the Llama 4 model family. It is designed to process both text and images and to generate multilingual text and code.
Maverick uses a Mixture-of-Experts (MoE) architecture with approximately 400 billion total parameters, while about 17 billion parameters are active for each token. The model contains 128 routed experts plus a shared expert and supports a documented context window of up to 1 million tokens.
Unlike a conventional dense 17B model, Llama 4 Maverick still requires infrastructure capable of storing and serving its much larger total parameter set. The sparse MoE architecture reduces the amount of computation used for each token, but it does not reduce the complete model to the hardware footprint of a 17B model.
Meta released Llama 4 Maverick on 5 April 2025 as one of the first Llama models combining native multimodality with a Mixture-of-Experts architecture. Official weights are available for independent deployment under the Llama 4 Community License Agreement.
In practical terms, Llama 4 Maverick is intended for workloads such as multimodal assistants, image understanding, document analysis, coding, multilingual generation, long-context processing and AI systems that require self-hosted or managed deployment options.
Meta Llama 4 model card ↗