Llama 4 Scout and Maverick (Behemoth previewed)
The first MoE Llamas are Scout (17B active/109B total, 10M context) and Maverick (17B/400B, 128 experts), while the 2T-param Behemoth was only previewed.
Native multimodality with early fusion, alternating dense/MoE layers, iRoPE long-context attention and FP8 training on 30T+ tokens. Reception was poor. Maverick's LMArena rank came from an unreleased 'experimental chat' variant, and independent tests of the released weights ranked far lower.
- Date
- Saturday, 5 April 2025
- Lab
- Meta
- Kind
- open-weights
- Access
- open weights (restricted license)
Figures
| Measure | Value | Measured by |
|---|---|---|
| LMArena Elo (Maverick experimental chat version) | 1417 the released Maverick ranked 32nd on LMArena days later (TechCrunch) | company |
| Scout context window | 10M tokens 17B active, 16 experts | company |
A benchmark controversy followed. LMArena said Meta's interpretation of its policy did not match expectations and changed its policy (TechCrunch 2025-04-11). In a Jan 2026 FT interview Yann LeCun said results were 'fudged a little bit' (via The Decoder 2026-01-03). Meta's own denial/response was not retrieved. Behemoth was never released.
Sources
- ai.meta.com/blog/llama-4-multimodal-intelligence/
- techcrunch.com/2025/04/06/metas-benchmarks-for-its-new-ai-models-are-a-bit-misleading
- techcrunch.com/2025/04/11/metas-vanilla-maverick-ai-model-ranks-below-rivals-on-a-popular-
- the-decoder.com/?p=30626
This record was partly confirmed: some claims could not be checked on 6 October 2026. How we check