Towards dialect-inclusive recognition in a low-resource language: are balanced corpora the answer?
Lonergan, Liam, Qian, Mengjie, Chiaráin, Neasa Ní, Gobl, Christer, Chasaide, Ailbhe Ní
–arXiv.org Artificial Intelligence
ASR systems are generally built for the spoken 'standard', and their performance declines for non-standard dialects/varieties. This is a problem for a language like Irish, where there is no single spoken standard, but rather three major dialects: Ulster (Ul), Connacht (Co) and Munster (Mu). As a diagnostic to quantify the effect of the speaker's dialect on recognition performance, 12 ASR systems were trained, firstly using baseline dialect-balanced training corpora, and then using modified versions of the baseline corpora, where dialect-specific materials were either subtracted or added. Results indicate that dialect-balanced corpora do not yield a similar performance across the dialects: the Ul dialect consistently underperforms, whereas Mu yields lowest WERs. There is a close relationship between Co and Mu dialects, but one that is not symmetrical. These results will guide future corpus collection and system building strategies to optimise for cross-dialect performance equity.
arXiv.org Artificial Intelligence
Jul-14-2023
- Country:
- Oceania > New Zealand (0.04)
- North America > United States (0.04)
- Europe
- United Kingdom > England
- Cambridgeshire > Cambridge (0.04)
- Ireland > Leinster
- County Dublin > Dublin (0.14)
- United Kingdom > England
- Genre:
- Research Report (0.82)
- Technology:
- Information Technology > Artificial Intelligence
- Machine Learning (0.71)
- Speech (0.70)
- Natural Language (0.69)
- Information Technology > Artificial Intelligence