Eleven v4 and v4 Turbo: Speech Models and Voice Control

ElevenLabs released both models on September 28. They support more than 90 languages; v4 adds voice direction, while Turbo targets lower-latency conversational agents.

One generation for creation and conversation

ElevenLabs launched Eleven v4 and the lower-latency Eleven v4 Turbo on September 28, 2026. Both became available in ElevenCreative, ElevenAgents and ElevenAPI. V4 targets produced speech such as narration, dialogue and dubbing; Turbo is tuned for real-time conversational agents where a delay before the first audible reply affects the conversation. [1][2]

The company said v4 uses a new architecture but did not publish enough technical detail in the announcement to explain its internals. The visible change for a user is more control over delivery: natural-language direction and inline audio tags can steer emotion, pronunciation and sound effects, while the model uses surrounding text and speaker context when generating dialogue. [1]

What the new models claim to improve

Both models support more than 90 languages, compared with 70 in the previous generation according to TechCrunch. ElevenLabs said the models preserve a voice's identity across longer passages and languages, and that instant cloning can start from 10 seconds of source audio. The release also adds Professional Voice Clone support to v4. [1][2]

Atlas interpretation: ElevenLabs reported a median time to first speech of about 150 milliseconds for Turbo in its WebSocket test and said blind listeners preferred v4 to named competing models in roughly three quarters of comparisons. The latency test excluded network time, and the preference test was run by ElevenLabs. Those numbers describe the company's test conditions; they do not establish the delay or voice preference a particular deployed agent will achieve. [1]

Speech quality and turn-taking in the same release

Atlas interpretation: For a producer, direction tags and context-aware multi-speaker speech make the model a tool for shaping a performance instead of merely reading text. For an agent builder, Turbo connects that speech generation to ElevenAgents with a shorter claimed wait for the first words. TechCrunch covered both use cases on launch day, while the official announcement confirms the models were already available through the company's products and API. [1][2]

Sources

  1. Introducing Eleven v4, our most emotive model

    ElevenLabs · Sep 28, 2026

  2. ElevenLabs’ new v4 speech model supports more expression control and 90 languages

    TechCrunch · Sep 28, 2026