← News

ElevenLabs Eleven v4 release graphic — Eleven V4 wordmark on green-to-orange motion blur

28 Sep 2026

ElevenLabs

ElevenLabs launches Eleven v4 and v4 Turbo speech models

ElevenLabs on Monday released Eleven v4 and its low-latency twin Eleven v4 Turbo, saying the new speech models add finer expression control, stronger voice cloning from as little as 10 seconds of audio, support for more than 90 languages, and median first-speech latency around 100–150 milliseconds for voice agents.

Voice agents and creative read-aloud tools have forced a choice between a voice that sounds alive and a voice that answers fast. ElevenLabs says v4 is one generation that does both, with finer control over delivery and a low-latency twin built for live agents.

On Monday, 28 September 2026, ElevenLabs launched Eleven v4 and Eleven v4 Turbo. The post is on the company blog, written by Mati Staniszewski and Piotr Dabkowski, and dated September 28, 2026. Both are text-to-speech models: software that reads written words out loud. ElevenLabs says they are available now in ElevenAgents, ElevenCreative, and through the ElevenAPI. ElevenAgents is the product for a voice agent, software that talks with a person while the conversation is still going. ElevenCreative is the creative studio. The ElevenAPI is the door other software uses to call these models. Those lines are ElevenLabs’s.

ElevenLabs calls v4 its most emotive text-to-speech model. Emotive means the voice is meant to carry feeling, not only the words. The company says the model is built to read tone, pacing, emotion, character, and context, and that it can sound dramatic, tender, urgent, comedic, or conversational while keeping the speaker’s identity. It also says conversations with more than one speaker sound more natural, because a voice answers what was just said instead of separate lines stuck together. Those descriptions are the company’s.

How a user steers a line. ElevenLabs says people can describe the delivery in ordinary language, and can drop inline audio tags into the script. A tag is a note in the text, such as [laughs], [said angrily in French accent], [light rain], or [phone buzzing]. The company says v4 follows those tags, and the written direction, more accurately than its earlier models. It also says support for the International Phonetic Alphabet is stronger. That alphabet is a standard way to spell a sound, so a name or a word can be pronounced on purpose. TechCrunch, reporting the same launch, wrote that ElevenLabs introduced these tags with its v3 model and is expanding them in v4, including letting a user stack several tags and have the model follow that sequence. The stacking line is TechCrunch’s. The examples and the “more accurately” line are on the company post.

Voice cloning is a copy of someone’s voice made from a recording. ElevenLabs says Instant Voice Clones can now capture a voice with high fidelity from about 10 seconds of audio. High fidelity means the copy stays close to the original. The company says the match to the source voice is significantly better, and that the same voice holds together across long generations, dialogue, narration, and lines that are generated again. Professional Voice Clones, the higher-fidelity option, are also supported on v4. Request stitching, chaining separate generations into one longer piece, is more reliable, which ElevenLabs says helps in its Studio and in the Reader app. Those claims are the company’s.

Both models support more than 90 languages, ElevenLabs says. A voice recorded in one language can speak the others, keep its identity, and take on an accent that sounds native to the new language. The company says that accent holds more firmly than before, so the voice does not slide back toward the original accent as a long clip plays. TechCrunch wrote that the previous version supported 70 languages, and that ElevenLabs said the biggest quality jump was in Japanese, Brazilian Portuguese, Mandarin, and Cantonese. The 70-language count and those four languages are TechCrunch’s account of what the company said. The blog post states the new total as more than 90.

Latency is the wait before speech starts. A millisecond is a thousandth of a second. ElevenLabs says strong voice models have usually been slower, so people chose between a voice that sounds alive and a voice that answers fast. For v4 Turbo, the company reports a median inference latency of about 100 milliseconds. Inference, here, is the model producing the sound. It says that is faster than the usual pause between two people talking. The same post reports a median time to first speech of about 150 milliseconds: the wait from the request until sound you can hear. The footnote says that 150-millisecond figure was measured in September 2026, with the network delay taken out, on v4 Turbo over a streaming connection, using the same scripts and default settings against Cartesia Sonic 3.6, xAI’s text-to-speech, Google Gemini Flash-Lite text-to-speech, and OpenAI’s GPT-4o mini text-to-speech. These are ElevenLabs’s measurements.

ElevenLabs says Artificial Analysis ranked v4 number one on its Provider Voice Arena leaderboard in September 2026. The company also says about 75 percent of listeners preferred v4 in blind tests, meaning the listener was not told which company made each clip. The footnote names the other models: Cartesia Sonic 3.6, Inworld TTS-2, Google Gemini 3.8 Flash-Lite text-to-speech, and Google Gemini 3.8 Flash text-to-speech. For each pair, graders heard the same line from v4 and from one competitor and judged which was more expressive and which sounded more natural. A tie counted as half. Those ranking and preference numbers are what ElevenLabs reports. They are not a rerun published by an outside lab in this announcement.

Where the fast model is aimed. ElevenLabs says v4 Turbo is built to work with ElevenAgents, and that its research and engineering teams tuned the model and that agent product together, rather than bolting a voice onto someone else’s agent. The post’s examples are a calm healthcare agent that has to pronounce medical words, and a fast-talking game character. TechCrunch reported that more than 55 percent of ElevenLabs’s business now comes from large companies, that the new model can start making audio as soon as the language model behind an agent starts writing the answer, and that it can treat a confrontation, an escalation, or a hold differently. A language model is the system that decides the words. Those three lines are TechCrunch’s.

TechCrunch’s same-day story also gives the business backdrop, which Monday’s blog post does not. It says ElevenLabs raised $500 million from Sequoia earlier this year, at an $11 billion valuation. A valuation is the price that round put on the whole company. It says annualized revenue, a recent sales pace stretched across a year, climbed from roughly $330 million at the start of 2026 to over $600 million, and that the company has more than 800 people. It says there are rumors of a later round that would value the company at $22 billion. It quotes Staniszewski, from a recent interview, aiming for a public stock listing “in the next years,” with no date. Those figures and that listing line are TechCrunch’s.

The picture is ElevenLabs’s launch graphic for this release. The words “Eleven V4” sit on a green-to-orange motion blur, with a line that calls the model the company’s fastest and most emotive voice model. It is the release image.

RELATED

ONLINE…

Comments

guidelines

Loading…

Loading…

Sources