The Next Leap in ASR for a Multilingual World We are releasing Saaras V4, the latest generation of our speech recognition model. This post covers the work behind Saaras V4, including the data used to train it, our pre-training approach, the architectural and decoding changes that support multiple output modes, and the evaluation methodology behind our results. We also report benchmark performance across Indian languages, English, and challenging real-world speech conditions. Highlights Saaras V4 uses an audio encoder, and an LLM decoder. The LLM decoder is a 3B hybrid state space model completely trained from scratch in-house Saaras V4 achieved SOTA performance on all 22 Indian languages and reaffirms our commitment to even the most low-resource Indian languages Saaras V4 achieves lowest WER on seven global English datasets, six being foreign English dialects and one being Indian English You can choose the format that works best for you: Saaras V4 supports five transcript formats: transcribe, translate, verbatim, translit, and codemix Saaras V4 is robust to noisy audios and code mixing. It also handles dialectical variation seamlessly …
Top stories
