Africa-Centric Self-Supervised Pre-Training for Multilingual Speech Representation in a Sub-Saharan Context
Caubrière, Antoine, Gauthier, Elodie
–arXiv.org Artificial Intelligence
We present the first self-supervised multilingual speech model trained exclusively on African speech. The model learned from nearly 60 000 hours of unlabeled speech segments in 21 languages and dialects spoken in sub-Saharan Africa. On the SSA subset of the FLEURS-102 dataset, our approach based on a HuBERT$_{base}$ (0.09B) architecture shows competitive results, for ASR downstream task, compared to the w2v-bert-51 (0.6B) pre-trained model proposed in the FLEURS benchmark, while being more efficient by using 7x less data and 6x less parameters. Furthermore, in the context of a LID downstream task, our approach outperforms FLEURS baselines accuracy by over 22\%.
arXiv.org Artificial Intelligence
Apr-22-2024
- Country:
- Africa > Sub-Saharan Africa (0.25)
- North America > United States (0.05)
- Europe
- France (0.04)
- United Kingdom > England
- Cambridgeshire > Cambridge (0.04)
- Italy > Tuscany
- Florence (0.04)
- Asia
- Indonesia > Bali (0.04)
- Middle East > UAE
- Abu Dhabi Emirate > Abu Dhabi (0.04)
- Genre:
- Research Report (0.84)
- Technology:
- Information Technology > Artificial Intelligence
- Natural Language (1.00)
- Machine Learning (1.00)
- Speech > Speech Recognition (0.97)
- Information Technology > Artificial Intelligence