# Sarvam AI — Indic DiarBench: A Joint Diarization-ASR Benchmark Dataset for Indian Languages

- Company: Sarvam AI (sarvam.ai)
- Announced: 2026-08-11T06:44:00+00:00
- Category: not stated
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://www.sarvam.ai/blogs/indic-diarbench/
- Record: https://forck.live/items/18384-indic-diarbench-a-joint-diarization-asr-benchmark-dataset-for-indian-languages
- Subject: Sarvam / Bulbul / Saaras

The first open benchmark dataset for joint ASR and speaker diarization across all 22 scheduled Indian languages. Speech recognition benchmark datasets have evolved significantly in recent years, but there is still a gap. Most major datasets, especially for Indian languages, contain audio where only one person speaks at a time. Real conversations look very different. Meetings, podcasts, debates, and customer-support calls often involve interruptions, rapid turn-taking, and overlapping speech. As a result, a model can perform well on a single-speaker dataset but struggle when people are speaking over each other. This becomes even more important as speech systems move towards joint speaker-attributed ASR , systems that need to determine both who spoke and what they said . Traditionally, these have been treated as two separate problems. Diarization identifies who spoke when, while ASR transcribes the speech. But evaluating them separately can hide how errors in one affect the other. A system may transcribe the words correctly but assign them to the wrong speaker, or produce poor transcriptions when speech becomes short, choppy, and overlapping. …

---

Record: https://forck.live/items/18384-indic-diarbench-a-joint-diarization-asr-benchmark-dataset-for-indian-languages
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
