# Together AI — How speech models fail where it matters the most and what to do about it

- Company: Together AI (together.ai)
- Announced: 2026-02-23T00:00:00+00:00
- Category: research-paper
- Subject: Inference platform
- Models affected: Whisper, Deepgram, Phi-4, Whisper-Large
- Source: https://www.together.ai/blog/how-speech-models-fail
- Record: https://forck.live/items/2378-how-speech-models-fail-where-it-matters-the-most-and-what-to-do-about-it

Research shows that speech models have a 39% error rate on street name transcription, with an 18% accuracy gap between non-English and English primary speakers. A synthetic data technique called cross-lingual style transfer reduces errors by up to 60% with fewer than 1,000 samples.

## Evidence

Verbatim from https://www.together.ai/blog/how-speech-models-fail:

> we demonstrate that voice recognition systems struggle to understand street name pronunciations when speakers have diverse linguistic backgrounds — with an average transcription error rate of 39% across 15 state-of-the-art models, and an 18% accuracy gap between non-English and English primary speakers. We show that a synthetic data generation technique called "cross-lingual style transfer" can reduce these errors by up to 60% relative to the base model, using fewer than 1,000 training samples.

---

Record: https://forck.live/items/2378-how-speech-models-fail-where-it-matters-the-most-and-what-to-do-about-it
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
