A benchmark that tells the agent nothing
ARC Prize released ARC-AGI-3 on March 25, 2026, describing it as the first fully interactive benchmark in the ARC-AGI series: hundreds of turn-based game environments and thousands of levels, each handcrafted by a team of human game designers. There are no instructions, no rules, and no stated goals given to the agent. [1]
To succeed, an agent has to explore an environment on its own, work out how it behaves, discover what counts as winning, and carry what it learned into harder levels of the same game. The benchmark page describes agents perceiving the environment and choosing actions without relying on natural-language instructions, and without any pre-loaded knowledge or hidden prompts about the goal. [2]
Atlas interpretation: The earlier ARC-AGI benchmarks presented a model with a handful of static grid examples and asked it to infer a transformation rule. ARC-AGI-3 keeps a model or agent inside a running environment across many turns, where the only feedback is what happens after an action. That shifts the test from pattern completion on a fixed input to figuring out, through play, what is even being asked. [1][2]
Humans at 100 percent, frontier models under 1
ARC Prize reported that humans score 100 percent on ARC-AGI-3 and that frontier AI scores 0.51 percent. The launch post frames the earlier ARC-AGI benchmarks as having predicted and tracked prior breakthroughs, from reasoning models to coding agents, and presents this new gap as pointing to what comes next: the difference between AI that follows instructions and AI that can genuinely explore, learn and adapt in an unfamiliar setting. [1]
Atlas interpretation: The launch material does not publish the human sample size, the list of frontier models tested, or the exact scoring procedure behind those two numbers; it states the result and points researchers to a technical paper for the methodology. A benchmark's first-day score, on games no model has ever seen, mainly says that today's frontier models have not been built or tuned for this particular kind of open-ended play. Whether that gap says something deeper about exploration and goal inference, as opposed to something that a purpose-built agent harness closes quickly, is exactly what the following months of competition entries were meant to test. [1]
A competition and a stage, not just a paper
The launch paired the benchmark with ARC Prize 2026, a competition offering more than $2,000,000 in prizes across two tracks: a new Kaggle-style competition to build agents that play ARC-AGI-3 games, and a guaranteed grand prize on the original ARC-AGI-2 format for the best open-source solution. ARC Prize shared the announcement publicly at a launch event at Y Combinator's San Francisco headquarters, with a fireside conversation between ARC-AGI creator François Chollet and OpenAI chief executive Sam Altman on measuring intelligence on the path to AGI. [1]
Atlas interpretation: Putting a sitting frontier-lab chief executive on stage for the launch of a benchmark that his own company's models score near zero on is itself a signal: ARC Prize is positioning ARC-AGI-3 as a shared industry reference point rather than a niche academic test, the way its predecessor became one people cited when arguing about progress toward general intelligence. [1]
Sources
- Announcing ARC-AGI-3
ARC Prize · Mar 25, 2026
- ARC-AGI-3
ARC Prize · Sep 8, 2026