One model, not a pipeline of three
OpenAI announced GPT-4o, the o for omni, on May 13, 2024. The model accepts any combination of text, audio, image and video as input and can generate text, audio and image as output, in a single network rather than separate models stitched together for transcription, reasoning and speech synthesis. [1]
The headline number was latency. GPT-4o can respond to audio input in as little as 232 milliseconds, averaging 320 milliseconds, close to human conversational response time. OpenAI's earlier Voice Mode, which chained separate transcription, language and text-to-speech models, averaged 2.8 seconds on GPT-3.5 and 5.4 seconds on GPT-4. [1]
The live demo at OpenAI's San Francisco event showed the model being interrupted mid-sentence and resuming, picking up vocal tone and emotion, and switching between singing and flat delivery on request. Vision handling was demonstrated on the same call, reading a handwritten equation and describing a screen. [2]
The pricing was the part that reached the most people
OpenAI put GPT-4o in ChatGPT's free tier the same day, with paying Plus and Team subscribers getting up to five times the message allowance. That put a frontier-class model in front of anyone with an account, not just subscribers. [1]
In the API, OpenAI described GPT-4o as twice as fast as GPT-4 Turbo, at half the price, with five times the rate limit. On text, reasoning and coding benchmarks it matched GPT-4 Turbo rather than beating it; the gains OpenAI reported were in multilingual tokenization, audio and vision. [1]
Atlas interpretation: That combination, flat or improved capability plus a large price and access cut, is why the launch read as a distribution move rather than a pure capability jump. A model that ties GPT-4 Turbo on reasoning benchmarks is not the story; making that tier free and cheaper to build on is. [1]
The demo shipped ahead of the feature
Not everything shown on stage was available on May 13. Text and image capabilities began rolling out that day; the real-time voice mode demonstrated in the presentation was scheduled for alpha access to Plus users within the following month, and a macOS desktop app launched immediately with a Windows version to follow later. [2]
OpenAI said the model went through external red-teaming with more than 70 experts across domains including social psychology, bias and misinformation, and that its post-mitigation risk rating landed at medium across the categories it evaluated under the company's own preparedness framework. [1]
Sources
- Hello GPT-4o
OpenAI · May 13, 2024
- OpenAI debuts GPT-4o 'omni' model now powering ChatGPT
TechCrunch · May 13, 2024