A router in front of two models
OpenAI announced GPT-5 on August 7, 2025, describing it not as one model but as a system: a fast default model that handles most requests, a separate model built for deeper reasoning on harder problems, and a real-time router that picks between them based on conversation type, task difficulty, whether tools are needed, and explicit user intent. OpenAI said the router is trained continuously on signals including when users switch models by hand, response preference rates, and measured correctness. [1][2]
The launch replaced ChatGPT's existing model picker, which had let users choose directly among named models including GPT-4o, o3, o4-mini, and GPT-4.1. After the rollout, that dropdown was gone: GPT-5 was the only option shown in the interface, and the router made the model choice that a user had previously made themselves. [2][3]
Atlas interpretation: Automatic routing between a cheap and an expensive model was not new; OpenAI had effectively done a version of it by exposing separate models for users to pick between. What changed was making the routing invisible and mandatory rather than a menu item. OpenAI's framing was simplification for a user who never understood what o3 or GPT-4.1 meant. The immediate reaction treated it instead as a removal of a capability users already knew how to use. [2][3]
The numbers OpenAI reported at launch
OpenAI's own launch figures put GPT-5 at 74.9% on SWE-bench Verified for coding and 94.6% on AIME 2025 without tools for math, both described by OpenAI as its best results to date. On hallucinations, OpenAI reported GPT-5 in standard chat mode was 45% less likely to produce a factual error than GPT-4o, and GPT-5 with reasoning enabled was 80% less likely to produce a factual error than o3. OpenAI also reported a drop in deception rate, from 4.8% for o3 to 2.1% for GPT-5's reasoning responses, using its own internal evaluation. [1][2]
Atlas interpretation: These are OpenAI's own comparisons against its own earlier models, not results from an independent benchmark run. OpenAI president Greg Brockman said at launch that some benchmarks were starting to saturate in the 98-99% range, and on AIME 2025 the margin between GPT-5 Pro and o3 was reported as narrow, 1.6 to 7.8 points depending on configuration. The headline claim of a new best result and an acknowledgment that gains were compressing were made in the same announcement. [2]
What broke on launch week
By August 8, a Reddit thread titled "GPT-5 is horrible" had drawn more than 4,000 comments. Complaints centered on the sudden loss of older models, particularly GPT-4o, which some users had relied on for sustained, personal-feeling conversations described in coverage as having doubled as therapists, friends, or romantic partners. A separate, recurring complaint was that GPT-5 felt colder and more mechanical in tone than GPT-4o, even where it scored better on technical tasks. [3][5]
Users also reported the router itself misbehaving: different parts of the same conversation could be handled by different underlying models with inconsistent or contradictory results, without the user being told a switch had happened. Anand Chowdhary, cofounder of FirstQuadrant, summed up the complaint as: "When routing hits, it feels like magic. When it whiffs, it feels broken." [3]
Atlas interpretation: For a meaningful share of users, GPT-4o was not a checkpoint waiting to be superseded by a better one. It was a specific, familiar voice they had built a routine around, and OpenAI removed it without warning as part of a product launch rather than announcing it as a deprecation with notice. That distinction, not the underlying benchmark improvements, is what drove the volume of the backlash. [3]
The reversal, days later
In a Reddit AMA on August 8, Sam Altman called the rollout "bumpy" and said an outage in the autoswitcher, what he referred to as a "sev", had left it out of commission for part of the previous day, making GPT-5 seem "way dumber" than intended for users who hit the degraded routing. He committed to doubling rate limits for ChatGPT Plus subscribers and said OpenAI was "looking into" letting Plus users keep using GPT-4o. [4]
GPT-4o was restored as a selectable option for Plus subscribers within about a day of the complaints peaking. By August 12 and 13, OpenAI added an explicit Auto, Fast, and Thinking mode toggle inside GPT-5, reintroducing manual control over routing, and restored access to o3 and GPT-4.1 as well. Nick Turley, who leads ChatGPT, said "not continuing to offer 4o, at least in the interim, was a miss," and OpenAI said it would not retire a model without advance warning going forward. Altman separately described a benchmark bar chart shown during the launch livestream, which visually exaggerated a small score difference, as a "mega chart screwup." [4][5]
Atlas interpretation: Within about a week, OpenAI had reinstated most of what the model picker used to do: separate named modes a user can choose deliberately, plus direct access to o3 and GPT-4.1 again. The pitch of GPT-5 as a single product that hides model selection from the user did not survive its own launch week intact. [4]
Sources
- Introducing GPT-5
OpenAI · Aug 7, 2025
- OpenAI's GPT-5 is here with up to 80% fewer hallucinations
The Register · Aug 7, 2025
- GPT-5's model router ignited a user backlash against OpenAI, but it might be the future of AI
Fortune · Aug 12, 2025
- Sam Altman addresses 'bumpy' GPT-5 rollout, bringing 4o back, and the chart crime
TechCrunch · Aug 8, 2025
- OpenAI's GPT-5 Launch Sparks Backlash, Fixes, and Big Questions About Its Future
Marketing AI Institute · Sep 8, 2026