What the thirty-hour claim actually measured
Anthropic's launch post said it had "observed it maintaining focus for more than 30 hours on complex, multi-step tasks." This was not a published benchmark with a name and a leaderboard; it was Anthropic's own account of internal testing, illustrated with an example of the model rebuilding Claude.ai's own web client over roughly 5.5 hours and about 3,000 tool calls. No methodology, task suite, or third party was named alongside the 30-hour figure itself, and no outside lab reported reproducing it. [1][5]
On benchmarks Anthropic did name, Sonnet 4.5 scored 77.2% on SWE-bench Verified, averaged over 10 trials with no test-time compute and a 200K token thinking budget, and 82.0% in a "high compute" configuration using parallel sampling and rejection selection. On OSWorld, which scores a model operating a real computer desktop rather than just writing code, it reached 61.4%, up from 42.2% for the original Claude Sonnet 4 four months earlier. Pricing stayed at $3 per million input tokens and $15 per million output tokens, unchanged from Sonnet 4. [1]
Atlas interpretation: All of this is vendor-reported. Anthropic ran its own model against its own chosen test suite under its own chosen configuration, chose which competitor scores to show next to it, and published the result as a company blog post rather than a peer-reviewed evaluation. That does not make the numbers false, but it means "30 hours" and "77.2%" describe what Anthropic says its own model did on Anthropic's own terms, not a figure an outside party independently confirmed. [1]
An escalation of a four-month-old pitch
The endurance framing did not start with Sonnet 4.5. When Anthropic launched Claude Opus 4 and Sonnet 4 in May 2025, it said Opus 4 could "work continuously for several hours" and cited a customer that had run it "independently for 7 hours with sustained performance." That launch reported SWE-bench Verified scores of 72.5% for Opus 4 and 72.7% for Sonnet 4, both without extended thinking enabled. [2]
Atlas interpretation: Both launches emphasized longer-running work, but their duration claims describe different observations: a customer's seven-hour run with Opus 4 in May and Anthropic's report of more than 30 hours of sustained focus with Sonnet 4.5 in September. They do not establish a measured four-fold improvement. The later launch also reported higher SWE-bench Verified scores, under the configurations stated for each release. Together, the announcements show Anthropic putting more emphasis on sustained tasks while leaving the duration comparison dependent on its reported conditions. [1][2]
Why a model launch came with a tool update
The same day, Anthropic shipped Claude Code 2.0. The concrete changes were a checkpoint system that "automatically saves your code state before each change," letting a developer rewind with Esc twice or a /rewind command; subagents, which delegate a specialized piece of work such as standing up a backend API while a separate agent builds the frontend; a refreshed terminal interface; and a native VS Code extension, in beta, that shows Claude's edits as inline diffs in a sidebar rather than only in a terminal pane. [3][5]
A separate same-day post covered a different layer: context editing and a memory tool for the Claude Developer Platform, the API most agent builders use rather than the Claude Code CLI specifically. Context editing automatically clears stale tool calls and results as a long-running agent approaches its context limit; the memory tool lets Claude read and write files outside the context window so information persists across turns. Anthropic reported that combining the two improved performance by 39% over baseline on an internal agentic search evaluation and cut token consumption by 84% in a 100-turn web search test. [4]
Atlas interpretation: None of these features raise a benchmark score directly. Checkpoints, subagents, context editing, and memory exist because a model that can run for 30 hours needs infrastructure to survive 30 hours: something has to save its work so a bad step is recoverable, something has to keep its context from filling up with stale tool output, and something has to let it remember what it already tried. Shipping the tooling on launch day made the endurance claim usable rather than just impressive; a model that can run unsupervised for a day and a half is not very useful if the only way to recover from a wrong turn is to start over. [3][4]
A crowded six weeks
Sonnet 4.5 landed on September 29, 2025, about seven weeks after OpenAI's GPT-5 on August 7, and one day before OpenAI's Sora 2 video model, which shipped September 30. Coverage of the Sonnet 4.5 launch noted the surrounding pace directly: one write-up observed that Anthropic had shipped three major releases between May and September 2025 alone, with Anthropic co-founder Jared Kaplan telling reporters at launch that the company would "very likely" ship one or two more before year end. [5][6]
Atlas interpretation: Anthropic's own launch materials compared Sonnet 4.5 against GPT-5 and Gemini 2.5 Pro on SWE-bench Verified, and secondary coverage described Sonnet 4.5 as ahead of both on that chart. That comparison chart was Anthropic's own, built on Anthropic's own test harness against competitor models it did not train, which is a different exercise from a competitor rerunning the same test on its own infrastructure and is exactly the kind of vendor-versus-vendor benchmark comparison that produces different numbers for the same nominal test depending on who ran it. [1][5]
Atlas interpretation: None of the contemporaneous coverage found for this piece explicitly framed the Sonnet 4.5 launch itself as evidence of release-cadence fatigue among buyers. In September 2025 the pattern was still being described neutrally, as competitive pace rather than as a burden: labs shipping in the same narrow window, each claiming a new best result on an overlapping set of benchmarks, without an independent referee to settle which claim was strongest. [5]
Sources
- Introducing Claude Sonnet 4.5
Anthropic · Sep 29, 2025
- Introducing Claude 4
Anthropic · May 22, 2025
- Enabling Claude Code to work more autonomously
Anthropic · Sep 29, 2025
- Managing context on the Claude Developer Platform
Anthropic · Sep 29, 2025
- Anthropic Unveils Claude Sonnet 4.5, Claims AI Coding Crown Bolstered by New Developer Tools
WinBuzzer · Sep 29, 2025
- Sora 2 is here
OpenAI · Sep 30, 2025