Learning a Dual-Mode Speech Recognition Model via Self-Pruning

Liu, Chunxi, Shangguan, Yuan, Yang, Haichuan, Shi, Yangyang, Krishnamoorthi, Raghuraman, Kalinli, Ozlem

arXiv.org Artificial Intelligence 

In models typically operate under more storage and computational constraints contrast, most non-streaming models run from the server with fewer - e.g., on embedded devices - than any server-side full-context constraints. Therefore, instead of developing equally sized encoders, models. Motivated by the recent progress in Omni-sparsity supernet it is preferable to jointly build a compact streaming model and a training, where multiple subnetworks are jointly optimized in one large non-streaming model for real-world ASR applications. We note single model, this work aims to jointly learn a compact sparse ondevice that even though a single encoder is shared for both modes, we can streaming ASR model, and a large dense server non-streaming substantially prune it into a featherweight, e.g., about 30M parameters model, in a single supernet. Next, we present that, performing supernet as a streaming model, and use the original copy as a performant nonstreaming training on both wav2vec 2.0 self-supervised learning and encoder. Given the recent progress made in neural network supervised ASR fine-tuning can not only substantially improve the pruning [18, 19, 20, 21], we can specify a target sparsity level during large non-streaming model as shown in prior works, and also be able model training, prune the model weights accordingly before inference, to improve the compact sparse streaming model.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found