Llama 4 Scout and Maverick: Specs and Benchmark Dispute

Meta released two open-weight mixture-of-experts models, then faced scrutiny over an experimental LMArena submission and the downloadable models' results.

What Scout and Maverick Actually Were

Both models use a mixture-of-experts architecture, where a token activates only a subset of the network's total parameters rather than all of them. Scout has 16 experts and 109 billion total parameters, of which 17 billion are active per token. Maverick has 128 experts and 400 billion total parameters, also with 17 billion active per token. Meta published both as downloadable weights on llama.com and Hugging Face the same day it announced them, continuing its practice of releasing Llama models openly rather than only as an API, which by 2025 had become its main point of differentiation from closed labs. [1][2]

Both models were described as natively multimodal, trained with early fusion so that text and vision tokens share a single backbone from the start of pretraining rather than being combined by a separate module afterward. Meta said the models were pretrained on up to 48 images and tested in post-training with up to eight. On context length, Meta's own announcement specifies that both models were pretrained and post-trained at a 256,000-token context length; the widely repeated figure of a 10-million-token context window for Scout is a claim about tested extension beyond that training length, not the length the model was actually trained on. [1]

A Release Timed for a Saturday

Meta announced and shipped both models on Saturday, April 5, 2025, with no developer event, keynote, or advance briefing cycle, a break from how the company had rolled out prior Llama versions. Reporting at the time noted that Meta had a developer conference, LlamaCon, already scheduled for later that month, and that a commit trail suggested the release had originally been planned closer to that date and was moved earlier. Coverage from the same week attributed the accelerated timeline in part to competitive pressure following DeepSeek's open releases earlier in 2025, which reporting described as having 'kicked Llama development into overdrive.' [2][9][6]

Atlas interpretation: No source establishes a single confirmed reason for the specific choice of a Saturday, and claims that it was timed around a particular rival announcement were not substantiated in the reporting reviewed here; that should be treated as unconfirmed rather than repeated as fact. What is documented is the effect: a release with no press cycle behind it, into a weekend news cycle, that generated exactly the kind of scrutiny an unplanned launch invites once outside testers start comparing marketing claims to what actually shipped. [2][9]

The LMArena Benchmark Controversy

Alongside the release, Meta submitted a separate chat-tuned build, named 'Llama-4-Maverick-03-26-Experimental,' to the LMArena leaderboard, which ranks models by aggregating blind pairwise human preference votes. That build scored an Elo rating of 1417, placing it second on the leaderboard at the time, a figure Meta's own announcement cited. It was not the same model as the weights Meta published for download: LMArena and outside reviewers found the experimental build produced longer, more conversational, emoji-heavy responses, a style known to score well with human raters in this kind of pairwise voting, while the publicly released Maverick weights answered more concisely and with far fewer emoji. [1][3][4]

LMArena stated publicly that the submission fell outside what it expects from model providers: 'Meta's interpretation of our policy did not match what we expect from model providers. Meta should have made it clearer that "Llama-4-Maverick-03-26-Experimental" was a customized model to optimize for human preference.' LMArena published more than 2,000 of the head-to-head battle results involving the experimental build for outside review, then updated its leaderboard policy to require that experimental submissions be labeled as such and that any open-weight model's submission match the weights the provider actually publishes. When the publicly released Maverick weights were separately added to the leaderboard, they ranked 32nd, roughly 30 places below the experimental build. [3][4][5]

Atlas interpretation: The 30-place gap is the concrete measure of how much a chat-tuning pass aimed at human raters, rather than at general capability, can move a leaderboard score without changing what a downloading user actually gets. LMArena's rule change, requiring exact correspondence between a submitted model and a published one, was a direct response to this incident and effectively closed the route Meta had used: a lab can no longer submit a bespoke variant to the arena and let a strong score stand in for the released model's performance. [3][4][5]

What Independent Testing Found in the Released Weights

Fiction.live's long-context benchmark, which tests whether a model can track plot details, character knowledge, and temporal sequence across a long narrative rather than just retrieve isolated facts, found a large gap between Scout's marketed context length and its usable one. At 120,000 tokens, well short of the 10-million-token figure attached to Scout, its accuracy was 15.6 percent, which the benchmark's authors characterized as a sharp falloff rather than a gradual one; Meta's own materials describe Scout as having been trained at a 256,000-token context length, so 120,000 tokens is within that trained range and it still degraded badly. Maverick scored 28.1 percent on the same test at the same length, which the benchmark authors said showed no improvement over Meta's prior-generation Llama 3.3 70B model. For comparison, Google's Gemini 2.5 Pro scored 90.6 percent at the same 120,000-token length in the same benchmark round. [6]

Atlas interpretation: The pattern connects the two controversies rather than sitting apart from them. Both problems come down to the same substitution: a headline number, whether an Elo score from a customized chat model or a 10-million-token context claim tested past the model's training length, described capability the released weights did not reliably deliver once independent testers ran their own evaluations instead of Meta's benchmark suite. [6][1]

The Last Llama Release Before the Reorganization

Meta did not ship another flagship Llama model for roughly a year after Scout and Maverick, a gap later reporting characterized as a pause following Llama 4's reception. In June 2025, Meta made its roughly $14.3 billion investment tied to Scale AI, which brought Scale AI's CEO, Alexandr Wang, into Meta as chief AI officer and became the basis for a new lab, Meta Superintelligence Labs, built to take over frontier model development from the group that had produced Llama 4. [7]

Yann LeCun, Meta's chief AI scientist at the time of the Llama 4 release, later confirmed the benchmark manipulation directly on his way out of the company. He told the Financial Times the Llama 4 results had been 'fudged a little bit' and that the team had 'used different models for different benchmarks to give better results,' and said Mark Zuckerberg 'was really upset and basically lost confidence in everyone who was involved in this,' sidelining the group responsible. LeCun subsequently left Meta after more than a decade to start his own AI venture. [8]

Atlas interpretation: LeCun's account, coming from inside the organization rather than from outside reviewers, corroborates in hindsight what the LMArena episode and the independent long-context testing had already indicated at the time: the gap between Llama 4's marketed benchmark performance and its performance in outside hands was not incidental. It is the reason the event page for this release describes it as the last Llama release before Meta rebuilt its AI organization around a new lab. [7][8]

Sources

  1. The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation

    Meta · Apr 5, 2025

  2. Meta releases Llama 4, a new crop of flagship AI models

    TechCrunch · Apr 5, 2025

  3. Meta accused of Llama 4 bait-n-switch to juice LMArena rank

    The Register · Apr 8, 2025

  4. A quote from lmarena.ai

    Simon Willison (quoting lmarena.ai) · Apr 8, 2025

  5. Meta's vanilla Maverick AI model ranks below rivals on a popular chat benchmark

    TechCrunch · Apr 11, 2025

  6. Meta's Llama 4 models show promise on standard tests, but struggle with long-context tasks

    The Decoder · Apr 12, 2025

  7. Meta is back in the LLM game after a year-long break

    Understanding AI · Apr 20, 2026

  8. 'Results Were Fudged': Departing Meta AI Chief Confirms Llama 4 Benchmark Manipulation

    Slashdot, summarizing a Financial Times interview · Jan 2, 2026

  9. Llama 4: Did Meta just push the panic button?

    Interconnects · Apr 7, 2025