From the source
Sarvam 1: A small yet powerful LLM for Indic languages Large language models have demonstrated remarkable capabilities across diverse tasks, yet their development has predominantly focused on English and other high-resource languages.
This English-centric approach has created a significant technological gap for the billions of speakers of Indian languages.
While there have been efforts to introduce Indic languages to popular LLMs through continued pretraining, ground-up multilingual efforts like BLOOM are rare.
The effectiveness of these models is also limited by poor token efficiency for Indic scripts and insufficient high-quality training data for these languages.
We introduce Sarvam-1, a 2-billion parameter language model specifically optimized for Indian languages.
Built from the ground up to support 10 major Indian languages alongside English, Sarvam-1 demonstrates that careful curation of training data can yield superior performance even with a relatively modest parameter count.
Our work addresses two critical challenges in Indic language modeling: …




