ARC Prize: ARC-AGI Benchmarks, Competitions, AI Reasoning

ARC Prize is the nonprofit behind ARC-AGI, from static visual puzzles to ARC-AGI-3’s instruction-free games, where initial results put frontier models near zero.

kindNonprofit
founded2025
Event history

A benchmark built to resist being beaten

ARC Prize is a nonprofit built around a single benchmark family, ARC-AGI, rather than a product or a model. François Chollet, a machine learning researcher, introduced the underlying idea in a 2019 paper, "On the Measure of Intelligence": a test made of small visual puzzles, each showing a handful of input-output grid examples, that a person can usually solve by inference but that pattern-matching systems trained on large datasets tend to fail, because each puzzle uses a novel rule rather than one seen in training. Chollet ran the first public competition on the idea, on Kaggle, in 2020; the winning entry scored only about 20 percent, which is the result the benchmark was designed to produce. [1]

Chollet and Mike Knoop co-founded ARC Prize to run the competition on an ongoing basis. It grew from Kaggle competitions into the ARCathon series, and in 2025 converted into a nonprofit foundation, with Greg Kamradt as president, "to advance the mission of guiding open-source AGI research." The 2025 competition carried a prize pool above $725,000 and introduced ARC-AGI-2, a harder version of the original test. [1][2]

From static puzzles to interactive agents

On March 25, 2026, ARC Prize launched ARC-AGI-3, which it describes as the first interactive reasoning benchmark built to evaluate agentic intelligence. Rather than presenting a static grid puzzle with a single correct answer, ARC-AGI-3 drops an agent into one of several hundred hand-built game-like environments with no written instructions and no stated goal or rules; the agent has to work out what it is doing, and whether it is succeeding, purely by acting and observing the result. ARC Prize describes the measure as skill-acquisition efficiency: how quickly an agent can build a working model of a new environment from experience alone, rather than from being told what to do in natural language. [3]

On the initial results ARC Prize published, frontier models scored close to zero, roughly half a percent, on ARC-AGI-3's environments, while human testers solved effectively all of them. That gap, several orders of magnitude wide, is the point of the design: it is meant to stay hard for current systems specifically because it removes language instructions and repeated pattern exposure, the two things large language models tend to lean on most. [3]

Why a benchmark organization is on the timeline

Atlas interpretation: ARC Prize does not build models; it builds the tests models are measured against, which is a different kind of relevance to AI history than a lab's or a vendor's. Every frontier model release on the timeline eventually gets compared to something, and a persistent complaint about AI benchmarks is that they get memorized or gamed once a model's training data absorbs enough examples resembling them. ARC-AGI's design, novel puzzles or environments each time, evaluators withheld from public release, is a direct answer to that complaint. ARC-AGI-3's near-zero scores for frontier models are also a data point in the other direction from most of the timeline: most events here describe a capability arriving, while this one describes a capability that, as of the check date, still has not. [2][3]

Sources

  1. ARC Prize - History

    ARC Prize · Aug 20, 2026

  2. ARC Prize - Mission

    ARC Prize · Aug 20, 2026

  3. Announcing ARC-AGI-3

    ARC Prize · Mar 25, 2026