Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Research shows that speech models have a 39% error rate on street name transcription, with an 18% accuracy gap between non-English and English primary speakers. A synthetic data technique called cross-lingual style transfer reduces errors by up to 60% with fewer than 1,000 samples.
From the source
we demonstrate that voice recognition systems struggle to understand street name pronunciations when speakers have diverse linguistic backgrounds — with an average transcription error rate of 39% across 15 state-of-the-art models, and an 18% accuracy gap between non-English and English primary speakers. We show that a synthetic data generation technique called "cross-lingual style transfer" can reduce these errors by up to 60% relative to the base model, using fewer than 1,000 training samples.
together.ai