Three numbers a lab could publish
Five days after its chief executive called for pacing the frontier, Anthropic proposed three measurements it says any frontier developer could publish regularly with a public method: how much of its AI research AI itself performs, how well the actions of its AI agents are overseen, and how much of its research compute goes to safety. It also said the measures could trigger stronger requirements, such as a fixed testing window before a new model is used for further AI research. [1]
Anthropic's own readings
Its R&D Automation Index grades tasks on Epoch AI's scale of automation levels, from no AI involvement to AI working with no human in the loop. As of August 2026, Anthropic says, Claude led 26 percent of its AI research work, up from under 1 percent in February, and the share where Claude at least collaborated was above 90 percent. No measured work was fully autonomous. The figures come from sampling 20 percent of staff in each research department each week in July, with Claude judging the tasks, which Anthropic flags as a limitation. [1]
On oversight, about 30,000 agents were doing research and engineering work at any one time on its most-used internal platform in August. A monitor blocked 0.002 percent of over a billion agent decisions that month, about 1 in 47,000, and flagged roughly 100,000 transcripts a week, of which about 50 reached human review. On compute, in one week of July about 6 percent of AI research compute went to safety work, and about 12 percent of the compute for AI-driven research, which Anthropic calls deliberately conservative estimates and a snapshot of one week, not a fixed allocation. [1]
A disclosure, not yet a standard
Atlas interpretation: The 26 percent figure is the one that travels, and it describes something the pacing essay only asserted: AI doing a large and fast-growing share of the work of building AI. It is also a self-assessment, scored by Claude, from one company. Anthropic says it plans to embed independent evaluators from several organizations but had not yet done so. [1]
Atlas interpretation: Contrast it with METR's time-horizon measurements, which estimate from outside the length of task frontier models can complete with 50 percent reliability. Anthropic's index measures, from inside, how much of one lab's own work its models already do. [1][2]
Sources
- Measurements for understanding the pace of AI development inside frontier labs
Anthropic · Sep 17, 2026
- Time Horizon 1.1
METR · Jan 29, 2026