Gemini 2.5 Pro: Launch Benchmarks & Experimental Access

Google's March 2025 launch made Gemini 2.5 Pro free in AI Studio as an experimental model; the page checks its reported benchmark scores and launch caveats.

What Google claimed at launch

Google DeepMind released Gemini 2.5 Pro on March 25, 2025, describing it in its own announcement as its "most intelligent AI model" to date. The initial release went out as an experimental build: free in Google AI Studio and available in the Gemini app to Gemini Advanced subscribers, with Vertex AI access and pricing both described only as coming later. [1]

Google reported that Gemini 2.5 Pro led GPQA Diamond, a set of graduate-level science questions, at 84.0 percent, and AIME 2025 competition mathematics at 86.7 percent. On Humanity's Last Exam it reported 18.8 percent without external tools, its best score on that benchmark without extra scaffolding. On SWE-Bench Verified, a suite of real GitHub issues, it reported 63.8 percent using a custom agent setup Google built specifically for the test. All four figures were Google's own reported numbers; none had been reproduced independently by a third party at launch. [1]

Gemini 2.5 Pro also debuted at number one on LMArena's leaderboard, which ranks models by human preference in blind pairwise comparisons. LMArena's own account reported it as the largest single score jump the leaderboard had recorded, roughly 40 points ahead of Grok 3 and GPT-4.5, landing at an Elo score of 1370. It had been tested there under the codename "nebula" before Google attached its name to the result, a routine practice on the leaderboard for models not yet publicly announced. The model shipped with a one million token context window, with two million described as coming soon. [2][1]

An experiment, not a release

"Experimental" in Google AI Studio meant a free tier with rate limits Google did not publish alongside the announcement, no billing, and no stated service-level commitment. The blog post explicitly deferred both Vertex AI availability and pricing to a later date: "coming to Vertex AI soon," with pricing to follow "in the coming weeks." Developers could try the model that day; they could not yet build a paid product on it, or run it on Google's enterprise cloud platform. [1]

Atlas interpretation: The gap between the headline and what actually shipped is easy to miss in coverage that repeats "released" without qualification. A benchmark-topping model announced on a Tuesday and a production-grade, billed, enterprise-available model are different deliverables, and Google's own post treats them as two separate steps rather than one event. Reading the launch as an unqualified general release overstates how much of the product existed on March 25. [1]

How clean was the margin

About a month after the launch, researchers from Cohere Labs, Stanford, Princeton and other institutions published "The Leaderboard Illusion," an analysis of roughly two million Chatbot Arena battles across 42 providers. It found that a handful of large providers, including Google, tested substantially more private, unreleased model variants on the Arena than smaller labs could, then published only the variant that scored best; the paper reported Google and OpenAI had each received close to a fifth of all Arena battle data individually, versus roughly 29.7 percent combined for 83 open-weight models from everyone else. [3]

Atlas interpretation: Gemini 2.5 Pro's own path to the leaderboard, tested privately as "nebula" before its public debut, is a specific instance of the exact mechanism the paper describes, not proof that a worse-scoring version was discarded along the way. There is no public record of what "nebula" scored against alternate internal builds, so the honest position is that the practice existed and Gemini 2.5 Pro used it, not that the published Elo score was inflated by it. The GPQA, AIME and Humanity's Last Exam numbers sit on separate footing: they came from Google's own model card, and no independent lab had reproduced them by the time of launch. Trade coverage that week, such as InfoQ's, reported Google's numbers and the LMArena rank without flagging either point. [3][4]

From Bard's stumble to "back in the race"

The context makes the March 2025 reception legible. In February 2023, Google's Bard chatbot gave a wrong answer in its own promotional demo, claiming the James Webb Space Telescope took the first images of a planet outside the solar system; it did not. Google's shares fell roughly 7 to 8 percent that day, erasing on the order of 100 billion dollars in market value. [5]

Ten months later, Google's Gemini 1.0 launch video, "Hands-on with Gemini," was reported to have been edited: interactions the video presented as live, real-time exchanges had in fact been assembled from still images and shortened, cherry-picked text prompts. Both incidents fed a narrative, repeated through most of 2024, that Google was behind in the race its own researchers had helped start with the transformer architecture. [6]

Atlas interpretation: Gemini 2.5 Pro's reception reads as a direct answer to that narrative, independent of whether every underlying number holds up to scrutiny. InfoQ's write-up that week put it plainly: Google's latest model showed the company was "not only in the race but intends to shape its outcome." A wrong exoplanet fact and a staged demo video are concrete, verifiable failures; a leaderboard rank achieved partly through a testing practice available mainly to large labs is a subtler kind of claim, and press coverage in March 2025 mostly did not draw that distinction. [4]

Atlas interpretation: Whatever the precision of the March figures, the trajectory held. Gemini 2.5 Pro was the model that reopened the argument that Google could lead rather than follow, and eight months later Gemini 3 shipped directly into Google Search on its release day, a distribution advantage no competitor could match. The comparison this piece can support is narrow: the same company that lost 100 billion dollars in value over one wrong sentence in 2023 was, by late 2025, shipping a frontier model straight into Search itself. [7][5]

Sources

  1. Gemini 2.5: Our most intelligent AI model

    Google · Mar 25, 2025

  2. BREAKING: Gemini 2.5 Pro is now #1 on the Arena leaderboard

    LMArena · Mar 25, 2025

  3. The Leaderboard Illusion

    arXiv (Cohere Labs, Stanford, Princeton and co-authors) · Apr 29, 2025

  4. Google Introduces Gemini 2.5 Pro with Improved Reasoning and Coding Capabilities

    InfoQ · Mar 28, 2025

  5. Google shares lose $100 billion after company's AI chatbot makes an error during demo

    CNN · Feb 8, 2023

  6. Google's best Gemini demo was faked

    TechCrunch · Dec 7, 2023

  7. Gemini 3: Introducing the latest Gemini AI model from Google

    Google · Nov 18, 2025