# Gradium — Time to First Audio: Measuring and Reducing TTS Latency in Voice Agents

- Company: Gradium (gradium.ai)
- Announced: 2026-03-24
- Category: not stated
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://gradium.ai/blog/time-to-first-audio
- Record: https://forck.live/items/18325-time-to-first-audio-measuring-and-reducing-tts-latency-in-voice-agents
- Subject: Gradium TTS / Phonon

In natural conversation, the gap between one person finishing a sentence and the other starting to respond averages around 200 milliseconds . For voice agents this is the target to match. When such an agent is built on a cascaded pipeline involving a speech-to-text model (STT), an LLM, and a text-to-speech model (TTS), every millisecond spent in one component is a millisecond unavailable to the others. This post covers how to measure TTS latency, why cascaded architectures make it especially critical, and how Gradium compares to ElevenLabs, Mistral, and OpenAI on the metric that matters most: Time to First Audio . What Is Time to First Audio (TTFA) in Text-to-Speech? The key metric for real-time TTS is Time to First Audio (TTFA) : the elapsed time from sending the request to receiving the first playable audio samples. At Gradium, TTFA is one of the core performance metrics we optimize for. TTFA is straightforward to define but easy to measure incorrectly, because most TTS APIs use streaming responses that begin transmitting data before the full utterance is synthesized. …

---

Record: https://forck.live/items/18325-time-to-first-audio-measuring-and-reducing-tts-latency-in-voice-agents
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
