# Together AI — How Together AI built the world’s fastest speech-to-text stack

- Company: Together AI (together.ai)
- Announced: 2026-05-29T00:00:00+00:00
- Category: capability-change
- Subject: Inference platform
- Models affected: NVIDIA Parakeet-TDT 0.6B v3, OpenAI Whisper Large v3
- Source: https://www.together.ai/blog/how-together-ai-built-the-worlds-fastest-speech-to-text-stack
- Record: https://forck.live/items/2344-how-together-ai-built-the-world-s-fastest-speech-to-text-stack

Together AI announced optimizations to their speech-to-text (ASR) stack that reduce latency and improve throughput. By using TensorRT with multi-profile engines for real audio shapes, moving the decoder loop entirely to the GPU with conditional CUDA graph nodes, and collapsing CPU process boundaries to eliminate redundant copies, they achieved the fastest ASR stack according to Artificial Analysis. The stack serves NVIDIA Parakeet-TDT 0.6B v3 and OpenAI Whisper Large v3, with Parakeet transcribing 20 hours of speech in under 10 seconds.

## Evidence

Verbatim from https://www.together.ai/blog/how-together-ai-built-the-worlds-fastest-speech-to-text-stack:

> Together’s ASR stack serves the two lowest-latency speech-to-text models ranked by Artificial Analysis: NVIDIA’s Parakeet-TDT 0.6B v3 and OpenAI’s Whisper Large v3. The faster of the two, NVIDIA Parakeet-TDT 0.6B v3, can transcribe roughly 20 hours of speech, about the runtime of the Harry Potter film franchise, in under 10 seconds.

---

Record: https://forck.live/items/2344-how-together-ai-built-the-world-s-fastest-speech-to-text-stack
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
