Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Hugging Face partners with the Indian Institute of Science (IISc) and ARTPARK to improve accessibility and usability of the Vaani dataset, an open-source multi-modal, multi-lingual dataset representing India's linguistic diversity. The dataset, launched in 2022, targets over 150,000 hours of speech and 15,000 hours of transcribed text from 1 million people across 773 districts. Phase 1 (80 districts) is open-sourced, and Phase 2 is underway. The dataset supports various AI tasks including speech recognition, language modeling, and multimodal LLM enhancement.
From the source
The Indian Institute of Science (IISc) and ARTPARK partner with Hugging Face to enable developers across the globe to access Vaani, India's most diverse open-source, multi-modal, multi-lingual dataset.
huggingface.co