A second knob on the machine
OpenAI released o1-preview and o1-mini on September 12, 2024, in ChatGPT and its API. The company describes the models as trained with large-scale reinforcement learning to reason using an internal chain of thought before producing a final answer. That chain of thought is generated as reasoning tokens, billed to the user as output but not shown by default. ChatGPT displays only a summarized version of it. [1][2]
Atlas interpretation: Every prior GPT release scaled the same thing: more parameters and more pretraining data, computed once and then queried cheaply. o1 adds a second cost that scales at answer time rather than training time, in how many reasoning tokens the model spends before it replies. The seconds of compute per response in the summary above describes that second cost, not a bigger version of the first one. [1][3]
What 74, 83 and 93 percent actually describe
On the 2024 AIME math competition, OpenAI reported o1 averaging 74 percent (11.1 of 15 problems) with one sample per problem. That rose to 83 percent (12.5 of 15) taking a majority vote across 64 sampled attempts, and to 93 percent (13.9 of 15) when 1,000 samples were reranked by a separately trained scoring model. The company also reported the 89th percentile on Codeforces competitive-programming problems and accuracy on the GPQA science benchmark exceeding PhD holders answering questions outside their own specialty. [3]
Atlas interpretation: The 74 percent figure is the one that traveled through coverage: a single answer, no retries. The 93 percent figure needed a thousand candidate solutions and a scoring model to pick among them, which is a different pipeline than a chat reply. Repeating 93 percent on AIME without the sampling budget behind it describes a research setup as if it were the deployed model. [3]
Reasoning applied to its own safety rules
OpenAI's system card describes training o1 to reason about the company's safety policy in context before answering, an approach it calls deliberative alignment. On the system card's own harder refusal-evaluation set, the fraction of unsafe completions correctly refused rose from 0.713 for GPT-4o to 0.934 for o1-preview, and OpenAI reported the largest jailbreak-robustness gains it had measured to that point. [2]
Why the reasoning stays hidden
OpenAI's stated reasons for hiding the chain of thought are competitive advantage and keeping a model's raw, sometimes policy-violating deliberation out of a user's hands. Simon Willison, writing the same day, called the decision "a big step backwards" for anyone trying to audit what the model actually did to reach an answer, and noted OpenAI researcher Jason Wei's own caveat that strong benchmark scores do not necessarily translate into something a user can feel in ordinary use. [4]
Sources
- Introducing OpenAI o1-preview
OpenAI · Sep 12, 2024
- OpenAI o1 System Card
OpenAI · Sep 12, 2024
- Learning to Reason with LLMs
OpenAI · Sep 12, 2024
- Notes on OpenAI's new o1 chain-of-thought models
Simon Willison · Sep 12, 2024