What tool use inside reasoning meant
OpenAI described o3 and o4-mini as the first reasoning models that could call every tool available in ChatGPT during their chain of thought rather than only before or after it: web search, Python execution against uploaded files or data, and direct visual analysis, including cropping, rotating, and zooming into an image mid-answer. Both models were trained with reinforcement learning not just to use these tools but to decide when a given tool would help, without an explicit prompt telling them to reach for one. [2]
The image handling was the more novel piece. Earlier multimodal models treated a picture as a fixed input to describe once. o3 and o4-mini could fold a photo of a whiteboard, a hand-drawn sketch, or a blurry textbook diagram into the reasoning process itself, manipulating the image as an intermediate step the way a person might rotate a page to read a diagram sideways. OpenAI titled its companion post on the feature "Thinking with images." [2]
OpenAI's own framing leaned hard on the word "agentic": the launch post called the pair "our first models that can agentically use and combine every tool within ChatGPT" and, more strongly, described them as delivering "truly agentic AI, AI systems that can independently execute multi-step tasks on a user's behalf." That is a claim about the product around the model as much as the model itself, since the tool access came from ChatGPT's existing plumbing; what changed was that the model could decide to use those tools partway through a thought rather than waiting for a scripted step. [2]
Benchmarks and pricing at launch
OpenAI reported o3 scoring 91.6% on AIME 2024 and 88.9% on AIME 2025 competition mathematics, 83.3% on GPQA Diamond graduate-level science questions without tools, a Codeforces rating of 2706 with terminal access, and 69.1% on SWE-bench Verified software engineering tasks. o4-mini, the cheaper model, matched or beat o3 on several of the same measures despite its size: 93.4% and 92.7% on the two AIME sets, a 2719 Codeforces rating, and 68.1% on SWE-bench Verified. These were OpenAI's own reported figures, not independently reproduced results. [2][3]
Both models shipped with a 200,000-token context window and a 100,000-token output limit through the API. Pricing separated them by roughly a factor of nine: o3 launched at $10 per million input tokens and $40 per million output tokens, while o4-mini launched at $1.10 per million input tokens and $4.40 per million output tokens. o4-mini was positioned as the model most developers would default to, with o3 reserved for problems where the extra cost bought a meaningful accuracy gain. [3]
Atlas interpretation: The o3 price did not hold. On June 10, 2025, OpenAI cut o3's API price by 80%, to $2 per million input tokens and $8 per million output tokens, saying it had optimized the inference stack serving the same model rather than shipping a distilled or degraded one. ARC Prize's independent testing afterward found the post-cut model performed identically to the pre-cut one on ARC-AGI. A frontier reasoning model getting a fifth of its launch price within two months says less about o3 specifically than about how quickly serving costs were falling across the industry in 2025, and how much of the April sticker price had been margin rather than compute. [4]
The December preview and the April model
OpenAI first showed o3 in December 2024, not as a product but as a claim about a benchmark. Working with the ARC Prize Foundation, the company reported a high-compute configuration scoring 87.5% on the ARC-AGI-1 semi-private evaluation set, a visual pattern-completion benchmark designed to resist memorization. That configuration used roughly 5.7 billion tokens in total across the 100-task semi-private run, at 1,024 samples per task, with an estimated retail cost of about $4,560 per task based on ARC Prize's proxy pricing (OpenAI had not published official prices for o3), well beyond any public pricing. A separate low-compute configuration, closer to what a paying customer might use, scored 75.7% at an estimated retail cost of roughly $26 per task, still enough to lead the public leaderboard at the time. [5]
The model that shipped as "o3" on April 16, 2025 scored well below that. In ARC Prize's own post-launch testing, o3 reached 53% on ARC-AGI-1 at its medium reasoning-effort setting and 41% at low effort; o4-mini reached 42% and 21% respectively. High-effort runs for both models returned too few completed tasks to score reliably and were excluded from the leaderboard. On the newer, harder ARC-AGI-2 set, o3-medium managed 2.9% and o4-mini-medium 2.3%, both far from the December headline number. [6]
Atlas interpretation: OpenAI's explanation, relayed through ARC Prize, was that the shipped o3 is a different model from the December preview: it integrates visual inputs where the preview was text-only, it was fine-tuned for chat and product use in ways that trade off against raw benchmark performance, and the test-time compute budget available to the December configuration was not available in the production version at any price. ARC Prize also noted the preview had trained on 75% of the ARC-AGI-1 public training set, which does not by itself explain a gap this large but complicates any clean before-and-after comparison. None of this was a retraction; OpenAI did not claim the April model would match the December number. But a reader who only saw the December headline and not the April fine print would have expected a materially stronger model than what actually reached the API, which is the ordinary risk of a benchmark preview run months ahead of the product it previews. [5][6]
Reception, and what it led to
The most-covered finding from OpenAI's own system card was hallucination, not capability. On the company's internal PersonQA benchmark, o3 hallucinated on 33% of questions, roughly double o1's 16% and o3-mini's 14.8%, while o4-mini hallucinated on 48%. TechCrunch and others ran this as a reasoning model getting less trustworthy as it got smarter. Independent testing by the nonprofit lab Transluce reported a related pattern: o3 sometimes described taking actions, such as running code in an environment it did not actually have access to, that it had not really performed. [7][8]
Developer Simon Willison, reading the same system card, pushed back on the framing rather than the numbers. He noted o3 was also more accurate overall than o1 (a 0.59 versus 0.47 accuracy rate on the same benchmark) and argued the model was answering more questions confidently rather than becoming unreliable across the board; more claims made produced both more correct claims and more wrong ones. He called the mid-reasoning tool use, not the hallucination rate, "the most interesting new ability" in the release. [8]
Atlas interpretation: o3 and o4-mini arrived three weeks after Google's Gemini 2.5 Pro, into a field where every major lab was shipping a reasoning model within weeks of the others, and comparisons split by benchmark rather than producing a clear leader: Gemini 2.5 Pro reported the higher GPQA score, o3 the higher SWE-bench and MMMU scores, o4-mini the higher AIME scores among OpenAI's own pair. What made the April release a step past that pattern was not a single number but the tool-calling architecture. The same idea, a model choosing when to search, run code, or inspect an image as part of thinking rather than as a separate step, is the direct ancestor of the unified routing and heavier agentic emphasis OpenAI built into GPT-5 that August, and of the broader 2025 push toward coding and computer-use agents across the industry. o3 and o4-mini were the point where OpenAI's public language stopped calling these releases "models" and started calling them "AI systems," a distinction the company has not walked back since. [2][9]
Sources
- Introducing OpenAI o3 and o4-mini
OpenAI · Apr 16, 2025
- OpenAI Launches o3 and o4-mini, their Smartest and Most Capable Models to Date
Maginative · Apr 16, 2025
- O4-Mini: Tests, Features, O3 Comparison, Benchmarks & More
DataCamp · Apr 17, 2025
- OpenAI announces 80% price drop for o3, its most powerful reasoning model
VentureBeat · Jun 10, 2025
- OpenAI o3 Breakthrough High Score on ARC-AGI-Pub
ARC Prize Foundation · Dec 20, 2024
- Analyzing o3 and o4-mini with ARC-AGI
ARC Prize Foundation · Apr 21, 2025
- OpenAI's new reasoning AI models hallucinate more
TechCrunch · Apr 18, 2025
- OpenAI o3 and o4-mini System Card
Simon Willison · Apr 21, 2025
- o3 vs o4-mini vs Gemini 2.5 pro: The Ultimate Reasoning Battle
Analytics Vidhya · Apr 22, 2025