A browser, not a whole computer
Gemini 2.5 Computer Use is built on Gemini 2.5 Pro's visual understanding and reasoning, post-trained for UI control and exposed through a new computer use tool in the Gemini API. Given a screenshot, the user's goal and a history of recent actions, it returns one function call representing the next UI action: a click, a keystroke, a scroll, a drag. Google's model card lists 13 supported actions and states the model has "not yet been optimized for OS-level control or mobile control," though it "demonstrates promise for mobile UI control tasks." [2][1]
Atlas interpretation: That is a narrower claim than Anthropic's or OpenAI's computer-use tools make. Anthropic's version, shipped as a raw API a year earlier, and OpenAI's Operator, both target a full desktop environment. Google confined this model to the browser and let the benchmark table carry the argument that a narrower target can still win. [2][3]
The numbers, and where they came from
On the model card's own comparison table, Gemini 2.5 Computer Use scored 69.0% on the official Online-Mind2Web leaderboard and 88.9% on WebVoyager's, ahead of OpenAI's computer-using agent model at 61.3% and 87.0% on those same leaderboards. Run through an identical Browserbase harness instead, the picture narrows but holds: Gemini scored 65.7% and 79.9% against Claude Sonnet 4.5's 55.0% and 71.4% and OpenAI's model's 44.3% and 61.0%. On AndroidWorld, measured by Google DeepMind itself, Gemini scored 69.7% against Claude Sonnet 4.5's 56.0% and Claude Sonnet 4's 62.1%; the card notes OpenAI's model "could not measure" there because Google had no API access to it. On OSWorld, a full-desktop-control benchmark, Gemini's own entry reads "OS control not yet supported." [2]
Atlas interpretation: The comparison is Google's to construct, not a neutral leaderboard: the official-leaderboard rows mix each vendor's own harness, the Browserbase rows run every model through one harness Google did not write, and the AndroidWorld row has no OpenAI result to compare against at all. The gap between Gemini's official and Browserbase scores on Online-Mind2Web, 69.0% against 65.7%, is small; OpenAI's gap on the same pair, 61.3% against 44.3%, is not, which is its own evidence that harness choice moves these numbers more than a few points. [2]
A day after OpenAI's Dev Day
Google published the model the day after OpenAI's annual Dev Day, at which OpenAI announced new ChatGPT apps and continued pushing its own ChatGPT Agent tool, described as able to "complete complex tasks on your behalf." The Verge's report noted Google's model, unlike ChatGPT Agent and Anthropic's computer-use tool, has access only to a browser rather than a full computer environment, and flagged Google's own claim that the model "outperforms leading alternatives on multiple web and mobile benchmarks." [3][1]
Guardrails, and what the model still gets wrong
Google's model card lists the same general limitations as the underlying Gemini 2.5 Pro: hallucination, and weak causal understanding, complex logical deduction and counterfactual reasoning. The model's knowledge cutoff is January 2025. Because the tool acts on live web pages, the card separately flags prompt injection, where adversarial instructions embedded in a page's content can be mistaken for the user's own, and "unintentional model failure modes" such as clicking the wrong button or filling the wrong form, which the card says can lead to failed tasks or data exfiltration. [2]
The stated mitigations require the model to ask for human confirmation before high-stakes actions such as completing a financial or retail payment, accessing health or financial records, sending communications, or modifying files, and the API's terms prohibit developers from bypassing that confirmation where it is required. An inference-time filtering system is also meant to block the model from reaching sites known to host illegal or dangerous content. [2]
Sources
- Introducing the Gemini 2.5 Computer Use model
Google · Oct 7, 2025
- Gemini 2.5 Computer Use Model Card
Google DeepMind · Oct 7, 2025
- Google's latest AI model uses a web browser like you do
The Verge · Oct 7, 2025