GPT-4o Launch: Voice Rollout, Vision, Pricing, and Access

OpenAI introduced GPT-4o for text, audio, image, and video input, with lower API pricing and free ChatGPT access; real-time voice was promised later.

One model, not a pipeline of three

OpenAI announced GPT-4o, the o for omni, on May 13, 2024. The model accepts any combination of text, audio, image and video as input and can generate text, audio and image as output, in a single network rather than separate models stitched together for transcription, reasoning and speech synthesis. [1]

The headline number was latency. GPT-4o can respond to audio input in as little as 232 milliseconds, averaging 320 milliseconds, close to human conversational response time. OpenAI's earlier Voice Mode, which chained separate transcription, language and text-to-speech models, averaged 2.8 seconds on GPT-3.5 and 5.4 seconds on GPT-4. [1]

The live demo at OpenAI's San Francisco event showed the model being interrupted mid-sentence and resuming, picking up vocal tone and emotion, and switching between singing and flat delivery on request. Vision handling was demonstrated on the same call, reading a handwritten equation and describing a screen. [2]

The pricing was the part that reached the most people

OpenAI put GPT-4o in ChatGPT's free tier the same day, with paying Plus and Team subscribers getting up to five times the message allowance. That put a frontier-class model in front of anyone with an account, not just subscribers. [1]

In the API, OpenAI described GPT-4o as twice as fast as GPT-4 Turbo, at half the price, with five times the rate limit. On text, reasoning and coding benchmarks it matched GPT-4 Turbo rather than beating it; the gains OpenAI reported were in multilingual tokenization, audio and vision. [1]

Atlas interpretation: That combination, flat or improved capability plus a large price and access cut, is why the launch read as a distribution move rather than a pure capability jump. A model that ties GPT-4 Turbo on reasoning benchmarks is not the story; making that tier free and cheaper to build on is. [1]

The demo shipped ahead of the feature

Not everything shown on stage was available on May 13. Text and image capabilities began rolling out that day; the real-time voice mode demonstrated in the presentation was scheduled for alpha access to Plus users within the following month, and a macOS desktop app launched immediately with a Windows version to follow later. [2]

OpenAI said the model went through external red-teaming with more than 70 experts across domains including social psychology, bias and misinformation, and that its post-mitigation risk rating landed at medium across the categories it evaluated under the company's own preparedness framework. [1]

Sources

  1. Hello GPT-4o

    OpenAI · May 13, 2024

  2. OpenAI debuts GPT-4o 'omni' model now powering ChatGPT

    TechCrunch · May 13, 2024