From the source
At Gradium, we develop cutting-edge audio language models designed to deliver natural, expressive, ultra-low latency voice interactions at scale and capable of performing any voice task from voice cloning to conversational AI.
We provide access to text-to-speech (TTS) and speech-to-text (STT) models through our API, and the models can run on a variety of NVIDIA GPUs, from L4 to H100, enabling flexible deployment to match diverse performance and scalability needs.
Streaming Audio Models for Real-Time Voice Applications As these models are used to build real-time voice agents, the two most important performance metrics are: Time To First Audio (TTFA): the time it takes from a user initiating a connection to the API to receiving the first meaningful audio packet.
Real-time factor (RTF): how much faster than real-time audio gets generated.
A RTF of 2x means that it takes 0.5s to generate one second of audio.
For interactive voice AI applications, it is important for the real-time factor to be above 1 so that there is no audio skip.
A higher RTF is useful for applications where lookahead on audio outputs is needed to control another model, for example generative lipsync.
…





