A voice layer that hands off the thinking
GPT-Live-1 is the full-duplex voice model OpenAI first put in ChatGPT in July: it listens and speaks at the same time instead of waiting for a turn to end. On September 10, 2026, OpenAI made it available in the API. The model handles the conversation itself and can delegate reasoning and tool calls to a backend text model, such as GPT-6 Astra or a third-party model, instead of chaining separate speech recognition, language and speech synthesis systems. [1][3]
The voice layer costs $0.05 a minute, billed by the second, and the backend model and tools are billed separately. OpenAI listed telephony support for phone agents, from restaurant reservations to customer support. [1][3]
Atlas interpretation: Splitting the talking from the thinking lets a developer choose the price of each: a cheap backend for routine calls, a frontier model for hard ones, with the same voice in front. It also means the per-minute price is a floor rather than the full cost of a call. [1][3]
OpenAI's numbers, and a rival five days later
OpenAI said GPT-Live-1 improved on its previous realtime voice model, GPT-Realtime-2.1, by 30 percentage points on Full Duplex Bench, and that paired with GPT-6 Astra at medium reasoning effort it ranked first on Tau3, a test of voice agents completing end-to-end tasks. Both are OpenAI's own evaluations. [1]
On Artificial Analysis's speech-to-speech leaderboard, GPT-Live-1 paired with Astra at medium effort scored 81.5, based on a single trial. Google's Gemini 3.8 Live Extended Thinking, released five days later, placed first at 82.6. [4]
Sources
- Build more natural voice experiences with GPT‑Live‑1 in the API
OpenAI · Sep 10, 2026
- OpenAI's GPT-Live-1 API lets developers build apps that talk and listen at the same time
The Decoder · Sep 10, 2026
- GPT-Live 1 Model | OpenAI API
OpenAI · Sep 22, 2026
- Speech to Speech Models and Providers Analysis
Artificial Analysis · Sep 22, 2026