Gemini 3.1 Flash Live: Audio Benchmarks, Preview, SynthID

Google released its streaming audio model in developer preview across Live products, reporting lower latency, speech benchmarks, and SynthID on every audio response.

What shipped, and where it runs

Google released Gemini 3.1 Flash Live in developer preview through the Gemini Live API in Google AI Studio, and put it behind three products at once: the consumer Gemini Live voice mode, Search Live, and Gemini Enterprise for customer experience. The model card lists it as an audio-to-audio model, and the API documentation gives it a 131,072 token input context and a 65,536 token output context, with text, image, audio and video accepted as input and text or audio returned. [1][2][3]

It connects over a WebSocket rather than a request and response call, streaming raw 16-bit PCM audio at 16kHz and, optionally, roughly one video frame per second. That streaming connection is what lets a caller interrupt the model mid-sentence, a behavior generally called barge-in, and what lets the model react to a pause or a change in tone before the speaker finishes a sentence. [3]

Atlas interpretation: The release follows Gemini 3.1 Pro by five weeks and reuses its version number, but the two are not the same kind of model. Flash Live has no text-reasoning benchmark attached to it anywhere in Google's materials; its numbers are all about holding up its end of a spoken conversation, which is a different job than the one Gemini 3.1 Pro was released to do. [1][4]

What the model actually improves on

Google's own claims are about listening rather than talking: better recognition of pitch and pace, including a user's frustration or confusion, and function calling that stays reliable across a multi-step tool sequence instead of degrading a few turns in. Google cites 90.8 percent on ComplexFuncBench Audio and 36.1 percent on Scale AI's Audio MultiChallenge with the model's thinking setting turned on, the second of which specifically scores how a model handles noisy or interrupted speech rather than clean dictation. [1][2]

Independent reporting describes the same change as a plumbing fix: earlier voice assistants ran speech through a separate transcription step, then a text model, then a separate speech synthesis step, and each handoff added delay and threw away acoustic information the transcript never captured. Gemini 3.1 Flash Live processes the audio stream directly instead of working from a transcript, which is what the pitch-and-pace claims and the lower latency both come from. [4]

Atlas interpretation: A developer-facing knob makes the tradeoff explicit: the API exposes a thinking level from minimal to high, defaulting to minimal for the fastest response, with higher settings trading latency for the reasoning depth behind the MultiChallenge number. That default matters more than the benchmark does for most deployments, since a voice agent tuned for speed will not be running the setting the benchmark used. [3][4]

What the SynthID watermark covers

Google says every audio response the model produces carries a SynthID watermark, described as imperceptible to a listener and interwoven into the output rather than added afterward, so it survives normal playback and lets a compatible detector flag the audio as machine-generated. [1]

Atlas interpretation: The watermark is a property of Google's audio, not a property of speech in general: it identifies output this model generated, not whether a given recording a listener encounters elsewhere was synthetic. Detecting it requires Google's own tooling, and neither the announcement nor the model card describes open, third-party verification. That is a narrower claim than "AI voices can be detected," and the distinction matters most in exactly the real-time phone and customer-service deployments this model is aimed at, where a listener has no independent way to check. [1][2]

Reach, and what preview status means

Google describes the model as inherently multilingual and says it supports conversations across more than 200 countries and territories, which follows from the model handling audio directly rather than through a separate translation or transcription stage tied to specific language packs. [1]

As of this check, the API documentation still lists the model as a preview release. Batch processing, asynchronous function calling, proactive audio and affective dialogue are all listed as unsupported, and Google has not published pricing for it. [2]

Sources

  1. Gemini 3.1 Flash Live: Making audio AI more natural and reliable

    Google · Mar 26, 2026

  2. Gemini 3.1 Flash live preview

    Google AI for Developers · Mar 26, 2026

  3. Gemini 3.1 Flash Audio (Flash Live, TTS) - Model Card

    Google DeepMind · Mar 26, 2026

  4. Google Releases Gemini 3.1 Flash Live: A Real-Time Multimodal Voice Model for Low-Latency Audio, Video, and Tool Use for AI Agents

    MarkTechPost · Mar 26, 2026