Exploring CTC Based End-to-End Techniques for Myanmar Speech Recognition
Chit, Khin Me Me, Lin, Laet Laet
–arXiv.org Artificial Intelligence
In this work, we explore a Connectionist Temporal Classification (CTC) based end-to-end Automatic Speech Recognition (ASR) model for the Myanmar language. A series of experiments is presented on the topology of the model in which the convolutional layers are added and dropped, different depths of bidirectional long short-term memory (BLSTM) layers are used and different label encoding methods are investigated. The experiments are carried out in low-resource scenarios using our recorded Myanmar speech corpus of nearly 26 hours. The best model achieves character error rate (CER) of 4.72% and syllable error rate (SER) of 12.38% on the test set.
arXiv.org Artificial Intelligence
May-13-2021
- Country:
- North America > United States
- New York (0.04)
- Europe
- Switzerland (0.04)
- Italy > Calabria
- Catanzaro Province > Catanzaro (0.05)
- Asia > Myanmar
- Yangon Region > Yangon (0.05)
- North America > United States
- Genre:
- Research Report (0.40)
- Technology: