From the source
Lead story
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
New text-to-speech model with emotional depth and low-latency variant
ElevenLabs launched Eleven v4, a text-to-speech model designed to interpret tone, pacing, emotion, character, and context, alongside a low-latency variant Eleven v4 Turbo with median inference latency of ~100ms.
Both models support over 90 languages, improved voice cloning with 10 seconds of audio, and more accurate audio tags and direction prompts.
The company states Eleven v4 was ranked #1 by Artificial Analysis and preferred by ~75% of listeners in blind tests over competing models.
From the source
Built on an entirely new architecture, Eleven v4 is our most emotive Text to Speech model ranked #1 by Artificial Analysis. Eleven v4 Turbo brings that same technology to low-latency use cases like agents. With a median inference latency of ~100ms, it can respond faster than the average pause between two people talking.
elevenlabs.io