SraVaani: Multilingual Speech Recognition Model

SraVaani

Context

Researchers at IISc’s SPIRE Lab, in collaboration with ARTPARK and Google, have developed SraVaani, an AI-based speech-recognition model designed to address India’s linguistic diversity.

About SraVaani

  1. SraVaani is an Automatic Speech Recognition (ASR) model that converts spoken language into written text.
  2. It supports 20 scheduled languages and 45 regional languages and dialects, including Garo, Angika, Chakma, Kokborok, Tulu, Bundeli and Bajjika.
  3. It can process speech across 10 scripts and automatically identify the language being spoken.
  4. The model uses a FastConformer-based architecture for efficient speech processing.

Project Vaani: Data Foundation

  1. Project Vaani, an IISc initiative, focuses on creating large-scale speech datasets representing India’s linguistic diversity, particularly languages with limited digital resources.
  2. It has collected about 31,255 hours of natural speech from more than 1.56 lakh speakers, covering 109 languages across 165 districts in 31 States/UTs.
  3. The dataset can support technologies such as Automatic Speech Recognition (ASR), Speech-to-Speech Translation (SST) and Natural Language Understanding (NLU).
  4. Its long-term goal is to create 1.5 lakh+ hours of speech data from around 10 lakh people across all 773 districts.

Significance

  1. Linguistic inclusion: Extends speech technology to languages and dialects with limited digital representation.
  2. Digital governance: Can improve access to e-governance and digital services in local languages.
  3. Public services: Potential applications include education, healthcare, banking and customer services.
  4. Indigenous AI: Strengthens India’s capacity to develop language technologies suited to its linguistic diversity.
  5. Language preservation: Digitising speech can support the documentation and preservation of India’s linguistic heritage.

Conclusion

SraVaani combines advanced speech-recognition technology with India’s expanding language-data resources. By supporting a wider range of Indian languages, it can contribute to inclusive digital services, indigenous AI development and linguistic preservation.

SraVaani FAQs

Q1. What is SraVaani?
Ans: SraVaani is a multilingual Indian speech-recognition model that converts speech into text and supports multiple Indian languages and scripts.

Q2. Who developed SraVaani?
Ans: It was developed by IISc’s SPIRE Lab in collaboration with ARTPARK and Google.

Q3. How many languages does SraVaani support?
Ans: It covers 20 scheduled languages and 45 regional languages and dialects.

Q4. What is the role of Project Vaani?
Ans: Project Vaani provided the linguistic foundation for SraVaani by collecting about 31,000 hours of speech data from over 1.56 lakh speakers across 165 regions.

Q5. What technology does SraVaani use?
Ans: It uses a FastConformer-based ASR architecture to recognise spoken language and convert it into text.