Qifusion-Net: Layer-adapted Stream/Non-stream Model for End-to-End Multi-Accent Speech Recognition
Chen, Jinming, Fang, Jingyi, Zheng, Yuanzhong, Wang, Yaoxuan, Fei, Haojun
–arXiv.org Artificial Intelligence
Others, accent identification (AID) models are used to generate embeddings [19, 20]. For instance, the Currently, end-to-end (E2E) speech recognition methods authors in [21, 22] suggest connecting accent embeddings and have achieved promising performance. However, auto speech acoustic features to adapt the acoustic model. In [23], they utilized recognition (ASR) models still face challenges in recognizing well-trained accent classifiers to extract accent embedding multi-accent speech accurately. We propose a layer-adapted fusion for layer-to-layer adaptation of E2E ASR models. A multi-task (LAF) model, called Qifusion-Net, which does not require framework was proposed in [21, 24] to jointly model ASR and any prior knowledge about the target accent. Based on dynamic AID tasks. All previous researches have significantly enhanced chunk strategy, our approach enables streaming decoding and the accuracy of accent speech recognition in specific contexts.
arXiv.org Artificial Intelligence
Jul-3-2024
- Country:
- Asia > China
- Beijing > Beijing (0.04)
- Guangdong Province (0.04)
- Asia > China
- Genre:
- Research Report (0.64)
- Industry:
- Education (0.34)
- Technology: