MUST: A Multilingual Student-Teacher Learning approach for low-resource speech recognition
Farooq, Muhammad Umar, Ahmad, Rehan, Hain, Thomas
–arXiv.org Artificial Intelligence
Student-teacher learning or knowledge distillation (KD) Student-teacher training or knowledge distillation (KD) has been previously used to address data scarcity issue for [13] has widely been used to distil the knowledge from either training of speech recognition (ASR) systems. However, a a single or multiple teacher models [14] to train a student limitation of KD training is that the student model classes model. This technique of transferring a teacher's knowledge must be a proper or improper subset of the teacher model to a student model either at output layer [13] or at intermediate classes. It prevents distillation from even acoustically similar stages [15] has been used for many tasks such as languages if the character sets are not same. In this work, the model compression [14, 16] and domain generalisation [17, aforementioned limitation is addressed by proposing a MUltilingual 18, 19]. The student model is trained with a combined objective Student-Teacher (MUST) learning which exploits a of minimising the KL-divergence loss for prediction of posteriors mapping approach. A pre-trained mapping model the teacher's posteriors (soft labels) and a classification loss is used to map posteriors from a teacher language to the student with the original training labels (hard labels).
arXiv.org Artificial Intelligence
Oct-28-2023
- Country:
- North America > United States (0.04)
- Europe > United Kingdom
- England > South Yorkshire > Sheffield (0.04)
- Genre:
- Research Report (0.50)
- Industry:
- Education (1.00)
- Technology: