Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Together AI announced optimizations to their speech-to-text (ASR) stack that reduce latency and improve throughput. By using TensorRT with multi-profile engines for real audio shapes, moving the decoder loop entirely to the GPU with conditional CUDA graph nodes, and collapsing CPU process boundaries to eliminate redundant copies, they achieved the fastest ASR stack according to Artificial Analysis. The stack serves NVIDIA Parakeet-TDT 0.6B v3 and OpenAI Whisper Large v3, with Parakeet transcribing 20 hours of speech in under 10 seconds.
From the source
Together’s ASR stack serves the two lowest-latency speech-to-text models ranked by Artificial Analysis: NVIDIA’s Parakeet-TDT 0.6B v3 and OpenAI’s Whisper Large v3. The faster of the two, NVIDIA Parakeet-TDT 0.6B v3, can transcribe roughly 20 hours of speech, about the runtime of the Harry Potter film franchise, in under 10 seconds.
together.ai