Claude Computer Use: API Beta, Benchmarks and Failure Rates

Anthropic gave Claude 3.5 Sonnet mouse and keyboard control through a public API beta. See its OSWorld scores, common failures and day-one risk controls.

A public beta, not a demo

Anthropic's announcement paired the capability with an upgraded Claude 3.5 Sonnet and previewed a smaller Claude 3.5 Haiku arriving later. Computer use was offered as a public beta API on Anthropic's own platform, Amazon Bedrock, and Google Cloud's Vertex AI: developers give Claude a screenshot of a desktop and it responds with mouse and keyboard actions, coordinates for where to click, text to type, keys to press, which get executed and fed back as the next screenshot. [1]

On OSWorld, a benchmark of real desktop and browser tasks, Claude 3.5 Sonnet scored 14.9 percent in a screenshot-only setting and 22.0 percent with extra steps allowed, against 7.8 percent for the next-best system Anthropic cited. Human performance on the same tasks runs 70 to 75 percent, a gap TechCrunch and Platformer both reported alongside the headline number rather than letting it stand alone. [1][3]

Slow, and wrong close to half the time

Anthropic's own language was blunt: computer use "remains slow and often error-prone" and struggles with scrolling, dragging, and zooming, and can miss a notification that only matters because of its timing. TechCrunch's testing of the airline-booking example Anthropic had used in its own demo found the model completed less than half of booking attempts and failed about a third of the time when told to start a return. [2]

Atlas interpretation: That an AI lab shipped a capability while describing it this way, in its own release notes, is the more interesting fact than the capability itself. Anthropic's advice to developers, start with low-risk tasks and add human checks for anything consequential, reads as a hedge against exactly the failure rates it had just published, and it is a rare instance of a vendor's stated limitations matching independent testing rather than undercutting it. [1][2]

The risks were named on day one

Anthropic said it retained screenshots from the beta for over 30 days and built classifiers meant to discourage the model from posting on social media, creating accounts, or interacting with government websites. Casey Newton's Platformer piece, published the same day, connected the retained-screenshot design to the backlash Microsoft's Recall feature had drawn months earlier over unencrypted local screenshots and default opt-in, and asked what businesses would need to know about what Anthropic did with the images of their screens. [2][3]

Atlas interpretation: The named risks, spam operations, fraudulent form-filling, AI-generated requests aimed at institutions that assume a human is on the other end, are downstream of the same trait that makes the tool useful: it acts on a screen the way a person would, so it inherits whatever a person could do with access to that screen, minus judgment. Anthropic's mitigations were procedural rather than technical, which is consistent with a capability its own numbers say is still mostly failing. [1]

Sources

  1. Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

    Anthropic · Oct 22, 2024

  2. Anthropic's new AI model can control your PC

    TechCrunch · Oct 22, 2024

  3. The AI agents have arrived

    Platformer · Oct 22, 2024