Goto

Collaborating Authors

 accent information


Qifusion-Net: Layer-adapted Stream/Non-stream Model for End-to-End Multi-Accent Speech Recognition

arXiv.org Artificial Intelligence

Others, accent identification (AID) models are used to generate embeddings [19, 20]. For instance, the Currently, end-to-end (E2E) speech recognition methods authors in [21, 22] suggest connecting accent embeddings and have achieved promising performance. However, auto speech acoustic features to adapt the acoustic model. In [23], they utilized recognition (ASR) models still face challenges in recognizing well-trained accent classifiers to extract accent embedding multi-accent speech accurately. We propose a layer-adapted fusion for layer-to-layer adaptation of E2E ASR models. A multi-task (LAF) model, called Qifusion-Net, which does not require framework was proposed in [21, 24] to jointly model ASR and any prior knowledge about the target accent. Based on dynamic AID tasks. All previous researches have significantly enhanced chunk strategy, our approach enables streaming decoding and the accuracy of accent speech recognition in specific contexts.


Accented Speech Recognition With Accent-specific Codebooks

arXiv.org Artificial Intelligence

Speech accents pose a significant challenge to state-of-the-art automatic speech recognition (ASR) systems. Degradation in performance across underrepresented accents is a severe deterrent to the inclusive adoption of ASR. In this work, we propose a novel accent adaptation approach for end-to-end ASR systems using cross-attention with a trainable set of codebooks. These learnable codebooks capture accent-specific information and are integrated within the ASR encoder layers. The model is trained on accented English speech, while the test data also contained accents which were not seen during training. On the Mozilla Common Voice multi-accented dataset, we show that our proposed approach yields significant performance gains not only on the seen English accents (up to $37\%$ relative improvement in word error rate) but also on the unseen accents (up to $5\%$ relative improvement in WER). Further, we illustrate benefits for a zero-shot transfer setup on the L2Artic dataset. We also compare the performance with other approaches based on accent adversarial training.