MUST: A Multilingual Student-Teacher Learning approach for low-resource speech recognition

Farooq, Muhammad Umar, Ahmad, Rehan, Hain, Thomas

arXiv.org Artificial Intelligence 

Student-teacher learning or knowledge distillation (KD) Student-teacher training or knowledge distillation (KD) has been previously used to address data scarcity issue for [13] has widely been used to distil the knowledge from either training of speech recognition (ASR) systems. However, a a single or multiple teacher models [14] to train a student limitation of KD training is that the student model classes model. This technique of transferring a teacher's knowledge must be a proper or improper subset of the teacher model to a student model either at output layer [13] or at intermediate classes. It prevents distillation from even acoustically similar stages [15] has been used for many tasks such as languages if the character sets are not same. In this work, the model compression [14, 16] and domain generalisation [17, aforementioned limitation is addressed by proposing a MUltilingual 18, 19]. The student model is trained with a combined objective Student-Teacher (MUST) learning which exploits a of minimising the KL-divergence loss for prediction of posteriors mapping approach. A pre-trained mapping model the teacher's posteriors (soft labels) and a classification loss is used to map posteriors from a teacher language to the student with the original training labels (hard labels).

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found