Unsupervised Domain Adaptation by Adversarial Learning for Robust Speech Recognition
Denisov, Pavel, Vu, Ngoc Thang, Font, Marc Ferras
–arXiv.org Artificial Intelligence
However, far-field speech, especially when recorded with single microphone, remains one of the major obstacles to achieving complete human parity, mainly because of challenging environments with a lot of noises and reverberations [3-7]. Another challenge is that it is almost impossible to collect data covering all recording environments to train and to test on due to variations of reverberations/noises and distances to microphones. While there were some advancements in this direction for a few widespread languages, a large number of low-resource languages will inevitably remain uncovered by such kind of resources. These facts motivate our interest for methods to improve robustness of acoustic model by utilizing available data, especially data from resource rich languages [8]. It is well known that a mismatch between training and testing conditions is likely to degrade accuracy of acoustic models.
arXiv.org Artificial Intelligence
Jul-30-2018
- Country:
- Europe > Germany > Baden-Württemberg
- Stuttgart Region > Stuttgart (0.04)
- Karlsruhe Region > Karlsruhe (0.04)
- Europe > Germany > Baden-Württemberg
- Genre:
- Research Report (1.00)
- Technology: