The model that had already shipped inside a product
OpenAI opened an improved version of Codex to a private API beta on August 10, 2021. Its own announcement described Codex as the model already powering GitHub Copilot, which OpenAI had built and launched with GitHub about a month earlier. What changed on this date was access: instead of suggestions inside a code editor, developers could now call Codex directly and build their own interfaces around it. [1]
OpenAI described Codex as a descendant of GPT-3, trained on natural language plus billions of lines of source code from publicly available sources, including public GitHub repositories. It was most capable in Python and also proficient in more than a dozen other languages, including JavaScript, Go, Perl, PHP, Ruby, Swift, TypeScript and Shell. OpenAI gave it a 14KB context window for Python code, compared with 4KB for GPT-3, and pointed to transpilation, code explanation and refactoring as uses it had already tried internally. [1]
Atlas interpretation: The distinction OpenAI drew was between a model that acts through the reader's mind and one that acts through working code. GPT-3 answers a prompt with text a person has to interpret and apply. Codex answers a prompt with code that can run against a real API, which is what let OpenAI frame it as a natural language interface to existing software rather than another writing tool. [1]
A free beta with the business model deferred
Access required an application describing the developer's intended project, reviewed by OpenAI's Codex team, with no published timeline or selection criteria. OpenAI offered the beta for free and said it planned to scale access as quickly as it judged safe, continuing the same incremental review approach it had used rolling out the GPT-3 API. Same-day coverage reported that OpenAI had not set a price or a date for turning the beta into a paid public API. [1][2]
OpenAI's CTO Greg Brockman and Codex lead Wojciech Zaremba framed the release around removing rote work rather than automating programming as a whole. Brockman called boilerplate and repetitive translation “the worst part of programming,” something Codex let developers do “with less keystrokes,” and Zaremba described the API as “a new way to interact with existing software.” [2]
Atlas interpretation: Free access paired with a deferred price is a common way to seed a new API surface. It let outside developers build the applications OpenAI could not anticipate itself, at the cost of leaving open, on launch day, exactly what running those applications would eventually cost. [1][2]
A benchmark score, and what it did and did not measure
The Codex paper, posted to arXiv about a month before this beta opened, introduced HumanEval, a set of standalone Python programming problems checked against unit tests rather than compared to reference text. A 12-billion-parameter Codex model solved 28.8% of those problems on a single attempt, against 0% for GPT-3 evaluated the same way; sampling 100 attempts per problem and keeping the best raised that to 70.2%. The paper distinguished this research model from the separate, further-tuned system deployed in Copilot. [3]
Atlas interpretation: Solving self-contained functions against unit tests is a narrower task than the one this beta was pitched for: building a natural language interface to arbitrary existing software. A one-attempt score on isolated problems says little about whether Codex's output would fit an application's own conventions, and the 70.2% figure describes picking the best of many generated candidates, not what a developer sees from a single request. [3]
The paper addressed the training data question directly. It argued that training on public data such as GitHub repositories had previously been treated as fair use, and reported that Codex rarely reproduced training data verbatim: exact matches to the training set occurred in well under 0.1% of generations it examined. It also noted that generated code is shaped by the user's own prompt and that the user retains full control over whether to accept or edit it. [3]
Atlas interpretation: That is OpenAI's own legal position stated in the paper backing this release, not an independent finding, and it is exactly the argument that later disputes over training on public code would contest. A near-zero verbatim-copying rate answers whether Codex regurgitates whole files; it does not settle whether training on code under restrictive open source licenses, or generating output that follows their patterns, requires the license terms those repositories carried. [3]
Sources
- OpenAI Codex
OpenAI · Aug 10, 2021
- OpenAI upgrades its natural language AI coder Codex and kicks off private beta
TechCrunch · Aug 10, 2021
- Evaluating Large Language Models Trained on Code
arXiv · Jul 7, 2021