# Gradium — Optimizing Quality vs. Latency in Real-Time Text-to-Speech AI Models

- Company: Gradium (gradium.ai)
- Announced: 2026-02-11
- Category: not stated
- Coverage: not counted
- Announcement: no
- Group: routine
- Source: https://gradium.ai/blog/optimizing-quality-vs-latency
- Record: https://forck.live/items/18327-optimizing-quality-vs-latency-in-real-time-text-to-speech-ai-models
- Subject: Gradium TTS / Phonon

At Gradium, we develop cutting-edge audio language models designed to deliver natural, expressive, ultra-low latency voice interactions at scale and capable of performing any voice task from voice cloning to conversational AI. We provide access to text-to-speech (TTS) and speech-to-text (STT) models through our API, and the models can run on a variety of NVIDIA GPUs, from L4 to H100, enabling flexible deployment to match diverse performance and scalability needs. Streaming Audio Models for Real-Time Voice Applications As these models are used to build real-time voice agents, the two most important performance metrics are: Time To First Audio (TTFA): the time it takes from a user initiating a connection to the API to receiving the first meaningful audio packet. Real-time factor (RTF): how much faster than real-time audio gets generated. A RTF of 2x means that it takes 0.5s to generate one second of audio. For interactive voice AI applications, it is important for the real-time factor to be above 1 so that there is no audio skip. A higher RTF is useful for applications where lookahead on audio outputs is needed to control another model, for example generative lipsync. …

---

Record: https://forck.live/items/18327-optimizing-quality-vs-latency-in-real-time-text-to-speech-ai-models
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
