A model that watches the screen instead of calling an API
Operator ran on a model OpenAI called Computer-Using Agent, or CUA, which combined GPT-4o's vision encoder with reasoning trained through reinforcement learning. It did not call a booking API or a structured tool interface. It took a screenshot of whatever was on screen, decided on a mouse click, a keystroke, or a scroll, executed that action in its own sandboxed browser, then took another screenshot and repeated the loop.This is the technical distinction that separates Operator from a typical API-calling agent: it operated the same graphical interface a person would, which meant it worked on sites that had never integrated with OpenAI, but also meant it was exposed to every bit of friction a human user hits, such as CAPTCHAs, unfamiliar layouts, and login screens. [1][2]
Atlas interpretation: The tradeoff was breadth for reliability. An API integration is fast and predictable but only exists where a partner built it. A screenshot-driven agent can in principle click through any website, at the cost of doing so far slower and more error-prone than the API path would be, since it has to parse the page visually instead of calling a function with typed arguments. [3]
Built to hand control back before it did anything it couldn't undo
OpenAI shipped Operator with three layers of restriction rather than letting it act freely. A "takeover mode" forced the user to type logins and payment details themselves rather than letting the model see or enter that data. A "watch mode" required the user to stay actively present while Operator worked on email or financial sites. And OpenAI had the model refuse certain categories of task outright, including banking transactions and other consequential decisions, rather than ask for confirmation on them at all.A separate monitoring system, described as combining automated and human review, was built to detect prompt injection attempts hidden in a page's content and to pause execution when it flagged suspicious behavior. Users could opt their browsing sessions out of model training and delete stored session data and site logins with one action. [1]
Atlas interpretation: That much scaffolding around a $200-a-month product is itself a signal. OpenAI was not confident enough in CUA's judgment to let it complete a purchase or send an email unsupervised, so it built manual checkpoints into the product to cover for that. In practice the checkpoints showed up as friction: an early hands-on account found Operator asked no clarifying questions about the user's location or existing accounts, required credentials to be re-entered by hand each session, and could not reach sites like YouTube or Reddit that blocked its crawler, making a six-item grocery order take roughly fifteen minutes. [3]
What the 38 percent figure is actually a score on
OSWorld is a benchmark of several hundred tasks run inside real Ubuntu, Windows, and macOS desktop environments, graded by execution rather than by a model judging the output: a task counts as passed only if the resulting file, setting, or application state actually matches what was asked for. It covers open-ended desktop work such as file management, office applications, and workflows that span more than one program, not just browser clicks. The benchmark's own reporting puts human performance at 72.36 percent of tasks completed.OpenAI's launch materials described CUA as setting new highs on WebArena and WebVoyager, two browser-only benchmarks, without publishing the underlying numbers on the announcement page itself. The specific figures that circulated afterward, and that match the OSWorld score already on record for this entry, were 38.1 percent on OSWorld and 58.1 percent on WebArena. [6][7]
Atlas interpretation: A 38 percent OSWorld score against a 72 percent human baseline means the model that shipped inside a $200-a-month product failed roughly three of every five tasks a person would have completed on a real desktop. WebArena's narrower, browser-only 58.1 percent looks better only because it excludes the file-system and cross-application work where CUA struggled most, which is why the January 2025 hands-on experience matched the benchmark gap rather than the marketing framing: reviewers found Operator slower, more expensive, and more frustrating than simply doing the task by hand. [3]
Gated to the tier OpenAI had just built for its reasoning model
The research preview launched January 23, 2025 at operator.chatgpt.com, restricted to ChatGPT Pro subscribers in the United States. OpenAI had introduced that $200-a-month Pro tier only weeks earlier alongside its o1 reasoning model, and Operator became the second flagship feature reserved for it. OpenAI said at launch that it planned to extend access to Plus, Team, and Enterprise subscribers and to additional countries, with the European rollout in particular running later than the US one. [1][2]
Absorbed into ChatGPT agent, then the standalone site went away
On July 17, 2025, OpenAI launched ChatGPT agent, describing it as combining Operator's ability to click through websites with Deep Research's ability to synthesize information across many sources and ChatGPT's conversational interface, available directly inside ChatGPT to Pro, Plus, and Team subscribers rather than at a separate URL. OpenAI's own release notes for that launch stated that "the standalone Operator experience at operator.chatgpt.com will be deprecated in the coming weeks," without committing to an exact date.Secondary reporting places the standalone site's actual shutdown at August 31, 2025, and separately reports that ChatGPT agent itself, the product that had absorbed Operator's browser control, was pulled from ChatGPT without an advance deprecation notice in early August 2026, with OpenAI's help center pointing remaining users toward other tools for browser-based work. [4][5][7]
Atlas interpretation: Operator's own life as a distinct, named product ran about six months, from research preview to a deprecation notice folding it into a broader agent surface. That the surface it was folded into was itself reportedly discontinued roughly a year after that says less about the underlying screenshot-driven technique, which persisted across both products, than about how quickly OpenAI has been willing to retire the wrapper around it. [7]
Sources
- Introducing Operator
OpenAI · Jan 23, 2025
- OpenAI launches Operator, an AI agent that performs tasks autonomously
TechCrunch · Jan 23, 2025
- OpenAI launches its agent
Platformer · Jan 23, 2025
- OpenAI launches a general purpose agent in ChatGPT
TechCrunch · Jul 17, 2025
- ChatGPT agent - release notes
OpenAI Help Center · Sep 8, 2026
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
OSWorld project (xlang.ai) · Sep 8, 2026
- OpenAI Operator
Wikipedia · Sep 8, 2026