From the source
BLEU vs. latency across all languages: Gradium ( s2s-translate ) versus gemini-3.5-live-translate and gpt-realtime-translate.
Today we are launching two models: stt-translate and s2s-translate . stt-translate collapses transcription and translation into a single step, so you speak in one language and it returns text in another directly, with no intermediate transcript to wait on. s2s-translate builds on it for a complete Speech-To-Speech experience: speech in one language goes in, and natural speech in another comes out.
Together they replace the usual three-model cascade (Speech-To-Text, Text-To-Text translation, Text-To-Speech) and deliver the combination of highest accuracy and lowest latency, with the ability to choose the target voice for the generated speech, including a clone of your own.
This post covers what stt-translate does, how we measure its quality, how s2s-translate pairs it with our TTS models for a full Speech-To-Speech workflow, and how both compare with gpt-realtime-translate and gemini-3.5-live-translate .
Real-time speech translation with Gradium s2s-translate .
What stt-translate does …





