ElevenLabs launched the Eleven v4 and Eleven v4 Turbo text-to-speech models in an update focused on improving vocal expressiveness, speaker identity consistency, and support for applications requiring near-instant responses. Both models are available through ElevenAgents, ElevenCreative, and ElevenAPI, and can be tried using free accounts.
Greater Control Over Tone and Context
The company says Eleven v4 is based on a new architecture capable of interpreting context, tone, speech rate, emotion, and personality when producing audio. This is intended to preserve the speaker’s identity in long texts while changing the delivery style according to the nature of the scene, such as dramatic, urgent, comedic, or everyday dialogue.
The model can also use previous sentences within a conversation to adjust the tone of subsequent sentences. ElevenLabs expanded the audio-tag system introduced with Eleven v3, making it possible to combine multiple instructions within a request and apply them in a specific order. Examples of the instructions include laughter, anger, a particular accent, light rain, and phone vibration. The company also announced improved support for the International Phonetic Alphabet (IPA) to help pronounce specific words and names.
Voice Cloning and Language Support
According to ElevenLabs, Eleven v4 can clone a voice from a recording no longer than 10 seconds. The company also made its Professional Voice Clone feature available again with the new model after it was not supported in Eleven v3.
Eleven v4 and Eleven v4 Turbo support more than 90 languages, compared with approximately 70 languages in Eleven v3. The company reports particular improvements in Japanese, Brazilian Portuguese, Mandarin, and Cantonese, as well as the ability of cloned voices to speak other languages with a more natural local accent, according to its evaluation.
What Changes in Practice with v4 Turbo?
Eleven v4 Turbo was designed for voice AI agents, call centers, and real-time applications directly affected by response time. The model supports bidirectional streaming, allowing audio generation to begin before the sentence has been produced in full.
The company’s tests indicate that the median time to first audio output is about 150 milliseconds, while the median inference time is approximately 100 milliseconds. These figures remain results from tests conducted by ElevenLabs and are not independent measurements. The company indicates that the model can begin speaking before the large language model has finished its response, and that it can also handle scenarios such as interruptions during a call, transferring to another representative, or placing the caller on hold.
Availability and Disclosed Limitations
Both models can be used through ElevenAgents, ElevenCreative, and ElevenAPI, and can be tried through free accounts. The free plan includes 10,000 credits per month, while paid plans start at $6 per month. Eleven v4 supports generating up to 10,000 characters in a single operation, with tools for connecting audio clips and maintaining rhythm and tone in long-form content such as audiobooks and podcasts.
The significance of the announcement lies in combining improvements in multilingual expressiveness with a separate path for low-latency response, but the actual quality of the results will remain linked to the language, the nature of the text, and response time in the usage environment, in addition to the need to verify whether voice cloning is appropriate for the intended use. The source did not provide independent tests directly comparing the two models with competitors.