Sound generated alongside the picture, not added after
Google introduced Veo 3 as its newest video generation model, capable of producing eight-second clips with synchronized sound effects and dialogue, a first for the company's video tools. Google described the model generating "traffic noises in the background of a city street scene, birds singing in a park, even dialogue between characters," and paired it with Flow, an online filmmaking tool combining Veo 3 with Google's Imagen 4 image generator and Gemini language model, letting users describe scenes in natural language and manage cast, locations and visual style in one interface. [1][2]
Both tools launched to US subscribers of Google AI Ultra, a $250 per month plan with 12,500 monthly credits. Each Veo 3 generation cost 150 credits, enough for about 83 videos before the plan's allotment ran out; extra credits cost 1 cent each in blocks of $25, $50 or $200, putting a single additional generation at roughly $1.50. [2]
A pipeline of three models, not one
Veo 3 is a system of separate components rather than a single model that jointly outputs picture and sound: a large language model interprets the text prompt, a video diffusion model generates the picture, and a distinct audio generation model applies sound to that video. The diffusion model itself works the way Stable Diffusion and similar image generators do, trained by adding noise to real video until it becomes static and learning to reverse that process, then starting generation from noise and a prompt and iteratively refining it. [2]
Atlas interpretation: "Generates the soundtrack with the video" describes what a viewer experiences, one generation producing a finished clip with sound already in it, more than what the architecture does internally. The audio still comes from a model applied to the video rather than one system solving both problems at once, which matters for anyone reading the result as evidence that video and audio synthesis have technically merged. [2]
One tell closes, others open
Veo 3 was not the first attempt at pairing AI audio with AI video. Meta had previewed a similar capability with Movie Gen the previous October, and Google DeepMind itself had shown an AI soundtrack-generating model as early as June 2024. What changed with Veo 3 was availability: dialogue and effects generated together with the picture, shipped as a consumer product rather than a research demo. [2]
Independent testing in the days after launch found consistent artifacts: on-screen subtitles that almost, but do not exactly, match the spoken words, an imitation of subtitled training footage rather than an accurate transcript; and in scenes with more than one person, dialogue that sometimes plays from the wrong character's mouth. Google said it embeds its SynthID watermark, designed to survive compression and editing, into every frame Veo 3 outputs, though this identifies a video as AI-generated only to someone who checks for it. [2]
Atlas interpretation: Silence closing as a giveaway does not mean AI video became undetectable, only that the giveaway moved. Garbled subtitles and misattributed dialogue are still artifacts a careful viewer can catch; they are simply different artifacts than the total absence of sound. What the launch changed most concretely was cost and access: a convincing eight-second scene with dialogue and sound, previously requiring specialized skill and editing software, became a $1.50 prompt available to anyone with a subscription. [2]
Sources
- Fuel your creativity with new generative media models and tools
Google · May 20, 2025
- AI video just took a startling leap in realism. Are we doomed?
Ars Technica · May 29, 2025