From the source
In natural conversation, the gap between one person finishing a sentence and the other starting to respond averages around 200 milliseconds.
For voice agents built on cascaded pipelines (STT, LLM, tool calls, TTS), TTS is the last stage before the user hears anything, and it inherits every upstream delay.
The latency budget for TTS itself is typically 200-300ms Time to First Audio (TTFA).
As of May 13, 2026, Gradium ranks first on Coval's TTS benchmarks across all latency metrics: P50 TTFA, latency range (P25-P75 spread).
This is not at the expense of quality metrics: Gradium provides state of the art Word Error Rate (WER).
Coval is an independent voice AI evaluation platform (YC S24); its TTS benchmark results are published at benchmarks.coval.ai/tts .
Gradium develops audio language models that power text-to-speech, speech-to-text, and voice cloning through a single API.
This post covers the methodology and the results of Gradium, ElevenLabs flash_v2.5 and Multilingual v2, Cartesia sonic-3, Rime arcana, and OpenAI.
What Coval is and why its TTS benchmarks matter …





