IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages
Javed, Tahir, Nawale, Janki Atul, George, Eldho Ittan, Joshi, Sakshi, Bhogale, Kaushal Santosh, Mehendale, Deovrat, Sethi, Ishvinder Virender, Ananthanarayanan, Aparna, Faquih, Hafsah, Palit, Pratiti, Ravishankar, Sneha, Sukumaran, Saranya, Panchagnula, Tripura, Murali, Sunjay, Gandhi, Kunal Sharad, R, Ambujavalli, M, Manickam K, Vaijayanthi, C Venkata, Karunganni, Krishnan Srinivasa Raghavan, Kumar, Pratyush, Khapra, Mitesh M
–arXiv.org Artificial Intelligence
We present INDICVOICES, a dataset of natural and spontaneous speech containing a total of 7348 hours of read (9%), extempore (74%) and conversational (17%) audio from 16237 speakers covering 145 Indian districts and 22 languages. Of these 7348 hours, 1639 hours have already been transcribed, with a median of 73 hours per language. Through this paper, we share our journey of capturing the cultural, linguistic and demographic diversity of India to create a one-of-its-kind inclusive and representative dataset. More specifically, we share an open-source blueprint for data collection at scale comprising of standardised protocols, centralised tools, a repository of engaging questions, prompts and conversation scenarios spanning multiple domains and topics of interest, quality control mechanisms, comprehensive transcription guidelines and transcription tools. We hope that this open source blueprint will serve as a comprehensive starter kit for data collection efforts in other multilingual regions of the world. Using INDICVOICES, we build IndicASR, the first ASR model to support all the 22 languages listed in the 8th schedule of the Constitution of India. All the data, tools, guidelines, models and other materials developed as a part of this work will be made publicly available
arXiv.org Artificial Intelligence
Mar-4-2024
- Country:
- Africa (0.04)
- South America > Suriname
- Marowijne District > Albina (0.04)
- North America
- Central America (0.04)
- United States
- Rhode Island (0.04)
- Maryland > Baltimore (0.04)
- New York > New York County
- New York City (0.04)
- Europe
- Greece (0.04)
- France > Provence-Alpes-Côte d'Azur
- Bouches-du-Rhône > Marseille (0.04)
- Asia
- Indonesia > Bali (0.04)
- Southeast Asia (0.04)
- Middle East > Saudi Arabia
- Asir Province > Abha (0.04)
- India
- West Bengal > Kolkata (0.04)
- Tripura (0.04)
- Tamil Nadu (0.04)
- Manipur > Imphal (0.04)
- Maharashtra > Mumbai (0.04)
- Gujarat > Gandhinagar (0.04)
- China > Shanghai
- Shanghai (0.04)
- Bangladesh > Dhaka Division
- Dhaka District > Dhaka (0.04)
- Genre:
- Research Report (0.81)
- Industry:
- Health & Medicine (1.00)
- Education (1.00)
- Consumer Products & Services (0.92)
- Transportation > Ground (0.67)
- Information Technology > Security & Privacy (0.67)
- Government > Regional Government
- Asia Government > India Government (0.66)
- Technology:
- Information Technology
- Communications > Social Media (0.93)
- Data Science (0.92)
- Artificial Intelligence
- Speech > Speech Recognition (1.00)
- Natural Language (1.00)
- Machine Learning (1.00)
- Information Technology