METR Time Horizon 1.1: New Tasks, New Estimates

METR expanded its task suite, moved to new evaluation infrastructure, and recalculated model time horizons. The update shows why a capability trend depends on its measurement design.

A duration at fixed reliability

METR's time horizon is not the time an AI system runs unattended. It is the estimated length of a task for a human professional at which an agent has a chosen probability of success. Its headline series uses the 50% point: researchers fit a success curve from task outcomes and human task times, then read off the duration at which predicted success is one half. [2]

The original work used multi-step software and reasoning tasks, and measured or estimated how long appropriately skilled humans took. It reported that task duration predicted agent success across that suite, but also said the result depends on choices such as the task set and the human comparison. It is a measurement of performance on that distribution of tasks, not a claim that an agent can reliably take over every project of the same calendar length. [2]

The revision changed what was being measured

TH1.1 expanded the suite from 170 to 228 tasks: 73 were added, 15 removed, and 53 updated. It more than doubled the number estimated to take people eight hours or longer, from 14 to 31, and moved the evaluation infrastructure from METR's Vivaria to the UK AI Security Institute's Inspect framework. METR recomputed estimates for 14 models rather than the 33 models in the earlier series. [1]

The larger suite narrowed some intervals at the frontier: METR gives Claude Opus 4.5's upper bound as 4.4 times its point estimate in TH1, versus 2.3 times in TH1.1. But those intervals remain wide, and only five of the 31 long tasks had human baseline times; the others used estimates. [1]

A revised trend is not a sudden capability jump

For models measured under both versions, METR says most changes came from the new task suite and run-to-run noise; the infrastructure change accounted for a relatively small share. The updated estimates generally stayed within the old confidence intervals. Since 2023, the fitted doubling time changed from 165 days under TH1 to 131 days under TH1.1, while the full-period stitched fit remained about 196 days. [1]

Atlas interpretation: That difference is evidence about the metric as well as the models. TH1.1's faster recent fit reflects older models moving down and newer ones moving up under a changed task distribution; it should not be read as a January 2026 step-change in model capability. The revision makes the benchmark more useful at longer task lengths, but also makes explicit that a time-horizon headline carries the selection rules and uncertainty of the evaluation behind it. [1][2]

Sources

  1. Time Horizon 1.1

    METR · Jan 29, 2026

  2. Measuring AI Ability to Complete Long Software Tasks

    METR · Mar 19, 2025