From the source
The first open benchmark dataset for joint ASR and speaker diarization across all 22 scheduled Indian languages.
Speech recognition benchmark datasets have evolved significantly in recent years, but there is still a gap.
Most major datasets, especially for Indian languages, contain audio where only one person speaks at a time.
Real conversations look very different.
Meetings, podcasts, debates, and customer-support calls often involve interruptions, rapid turn-taking, and overlapping speech.
As a result, a model can perform well on a single-speaker dataset but struggle when people are speaking over each other.
This becomes even more important as speech systems move towards joint speaker-attributed ASR , systems that need to determine both who spoke and what they said .
Traditionally, these have been treated as two separate problems.
Diarization identifies who spoke when, while ASR transcribes the speech.
But evaluating them separately can hide how errors in one affect the other.
A system may transcribe the words correctly but assign them to the wrong speaker, or produce poor transcriptions when speech becomes short, choppy, and overlapping.
…




