A coding model lands next to a flagship, again
OpenAI released GPT-5.3-Codex on February 5, 2026, the same day Anthropic shipped Claude Opus 4.6. OpenAI calls GPT-5.3-Codex its most capable agentic coding model to date, combining coding and reasoning in one model that runs about 25 percent faster than GPT-5.2-Codex, its immediate predecessor. [1]
Atlas interpretation: The two companies have landed a flagship on the same calendar day before: Claude's public launch and GPT-4 both shipped on March 14, 2023. Three years and a widened field of competing labs later, a second same-day landing is a data point for competitive timing pressure between the two, not yet a pattern this page can call routine. [1]
What moved, what did not, and two firsts
OpenAI's own comparison against GPT-5.2-Codex, both run at its highest reasoning effort setting: Terminal-Bench 2.0 rose from 64.0 percent to 77.3 percent, OSWorld-Verified from 38.2 percent to 64.7 percent, Cybersecurity Capture The Flag challenges from 67.4 percent to 77.6 percent, and SWE-Lancer IC Diamond from 76.0 percent to 81.4 percent. SWE-Bench Pro moved less, from 56.4 percent to 56.8 percent, and GDPval held flat at 70.9 percent, though OpenAI's table compares that figure at a lower reasoning-effort setting for the prior model than for GPT-5.3-Codex. [1]
OpenAI describes GPT-5.3-Codex as "the first we've directly trained to identify software vulnerabilities," and separately as "our first model that was instrumental in creating itself," saying it helped OpenAI teams debug the model's own training and deployment. [1]
Atlas interpretation: The largest jumps are on agentic and computer-use tasks (OSWorld-Verified nearly doubling, Terminal-Bench 2.0 up 13 points) rather than on SWE-Bench Pro, the benchmark closest to ordinary pull-request-sized coding work, which barely moved. That split matches OpenAI's own framing of the release as extending what the model can do around code, long sessions, tool use, computer control, rather than making it meaningfully better at the coding tasks the prior model already handled. [1]
A week later, a faster version to feel the difference
On February 12, OpenAI shipped GPT-5.3-Codex-Spark, a smaller, text-only, 128,000-token-context version of the model built for real-time coding, running on Cerebras's Wafer-Scale Engine hardware, the first release under a new OpenAI-Cerebras partnership. Both companies put its throughput at more than 1,000 tokens per second per user. [2][3]
OpenAI says the Cerebras-backed serving path cuts per-roundtrip overhead by 80 percent, per-token overhead by 30 percent, and time to first token by half, on tasks it says Codex-Spark completes in a fraction of the time GPT-5.3-Codex takes. It launched as a research preview for ChatGPT Pro users in the Codex app, CLI and VS Code extension, plus a small group of API design partners under separate, fluctuating rate limits. [2]
Atlas interpretation: Wafer-Scale Engine hardware is built for inference latency, not for training the largest models, so pairing it with a shrunk, text-only Codex variant rather than the full GPT-5.3-Codex is a fit between what the chip is good at and what a coding assistant needs most: fast turnaround on a tool call, not maximum context or multimodal input. The tradeoffs, a smaller model and a narrower context window, are the price of the speed rather than a side effect of it. [2][3]
Sources
- Introducing GPT-5.3-Codex
OpenAI · Feb 5, 2026
- Introducing GPT-5.3-Codex-Spark
OpenAI · Feb 12, 2026
- Introducing OpenAI GPT-5.3-Codex-Spark Powered by Cerebras
Cerebras · Feb 12, 2026