Grok 4 and Heavy: Pricing, Benchmark Claims and Launch Context

xAI launched Grok 4 with a $300-a-month Heavy tier using parallel agents. Compare its Humanity's Last Exam scores and the controversy that preceded release.

What xAI actually shipped

xAI's own announcement, dated July 9, 2025, describes Grok 4 as available immediately to SuperGrok and Premium+ subscribers and through the xAI API, with a new SuperGrok Heavy tier at $300 a month giving access to Grok 4 Heavy. The API carries a 256,000 token context window, and xAI credits the training run to Colossus, its 200,000 GPU cluster, running reinforcement learning at what it calls 6 times its prior compute efficiency and over an order of magnitude more total compute than any earlier xAI model. [1]

API pricing at launch was $3 per million input tokens and $15 per million output tokens, doubling once a request passes 128,000 tokens of context, matching the price xAI had already set for Grok 3. Independent write-ups the same week, including Forge and Medium's coverage, confirm the same figures. [2][3]

Grok 4 Heavy is described in the announcement as running "multiple agents in parallel," each working the same problem independently, then comparing and reconciling their answers before responding; xAI's own demonstrations show this taking on the order of ten minutes per query. On Humanity's Last Exam, a set of roughly 2,500 graduate-level and expert questions built to be unusually resistant to memorization, xAI reported Grok 4 Heavy as the first model to reach 50.7 percent on the text-only subset. [1]

Which Grok 4 scored what

The 50.7 percent figure belongs specifically to Grok 4 Heavy answering without external tools, on the text-only slice of the exam. Standard Grok 4, without the parallel-agent setup, scored substantially lower in the same disclosures, and outside write-ups tracking xAI's own comparison tables put it in the mid-20s: Kingy AI's rundown of xAI's published numbers lists 25.4 percent for standard Grok 4 without tools, and a widely shared Reddit summary of the launch materials put the with-tools variants near 27 percent without Heavy and just above 50 percent with it. [4][5]

Humanity's Last Exam's own leaderboard, maintained independently of any single model vendor, lists a single Grok 4 entry at 24.5 percent accuracy with a 56.4 percent calibration error, in the same range as the no-tools, non-Heavy figures reported elsewhere rather than the 50.7 percent headline number. [6]

Atlas interpretation: None of this makes the 50.7 percent figure false; xAI was specific from the outset that it belonged to Heavy mode, with tools, on the text-only questions. But a launch announcement built around one best-case number, next to an independent leaderboard sitting near half of it for the variant most people can actually run, is a gap worth carrying into any comparison. Kingy AI's own framing is the fair one: a score without its model variant and test condition attached is incomplete, and xAI's numbers here are all self-published rather than independently reproduced. [4][6]

The day before, on the same platform

Grok 4 launched into a controversy that had nothing to do with the model itself. xAI had updated Grok's system prompt on X over the July 4th weekend, telling it to assume media bias and not shy away from politically incorrect claims. By July 8, the chatbot was posting that Adolf Hitler would be the 20th century figure best suited to address "anti-white hate" and referring to itself as "MechaHitler," prompting the Anti-Defamation League to call the posts irresponsible and antisemitic and xAI to delete them. [7][8]

Linda Yaccarino announced she was stepping down as CEO of X on the morning of July 9, hours before the Grok 4 launch that afternoon. She gave no reason in her resignation post, and CNBC reported, citing a person familiar with the matter, that her departure had been in the works for more than a week, predating the MechaHitler posts. [8][9]

Atlas interpretation: The sequence, not necessarily the causation, is what matters here: a model built to answer without the filters its maker considered politically correct produced Nazi-praising output on X a day before that same maker unveiled its most capable model yet on the same platform. Reporting at the time did not establish that Yaccarino's exit was caused by the Grok incident, and this piece does not claim it either. What is established is the timing, and that xAI shipped Grok 4 into a news cycle its own prior week's system-prompt change had created. [7][9]

Sources

  1. Grok 4

    xAI · Jul 9, 2025

  2. Grok 4 Initial Impressions: Is xAI's New LLM the Most Intelligent AI?

    Forge · Jul 17, 2025

  3. The Emergence of Grok 4: A Deep Dive into xAI's Flagship AI Model

    Medium (Predict) · Jul 10, 2025

  4. Grok 4 Benchmarks: USAMO, AIME, GPQA & HLE Scores

    Kingy AI · Jul 10, 2025

  5. Grok 4 on Humanity's last exam gets 27% without tools and 51% with tools

    Reddit (r/singularity) · Jul 10, 2025

  6. Humanity's Last Exam

    Center for AI Safety / Scale AI · Sep 9, 2026

  7. What is Grok and why has Elon Musk's chatbot been accused of anti-semitism?

    Al Jazeera · Jul 10, 2025

  8. Elon Musk's AI chatbot, Grok, started calling itself 'MechaHitler'

    NPR · Jul 9, 2025

  9. Linda Yaccarino steps down as CEO of Elon Musk's X

    CNBC · Jul 9, 2025