On the Robustness of Arabic Speech Dialect Identification
Sullivan, Peter, Elmadany, AbdelRahim, Abdul-Mageed, Muhammad
–arXiv.org Artificial Intelligence
Arabic dialect identification (ADI) tools are an important part of the large-scale data collection pipelines necessary for training speech recognition models. As these pipelines require application of ADI tools to potentially out-of-domain data, we aim to investigate how vulnerable the tools may be to this domain shift. With self-supervised learning (SSL) models as a starting point, we evaluate transfer learning and direct classification from SSL features. We undertake our evaluation under rich conditions, with a goal to develop ADI systems from pretrained models and ultimately evaluate performance on newly collected data. In order to understand what factors contribute to model decisions, we carry out a careful human study of a subset of our data. Our analysis confirms that domain shift is a major challenge for ADI lected data to probe the limits of our transfer learning methods models. We also find that while self-training does alleviate this in a realistic data pipeline.
arXiv.org Artificial Intelligence
Jun-1-2023
- Country:
- North America > Canada
- British Columbia (0.04)
- Europe
- Spain > Catalonia
- Barcelona Province > Barcelona (0.04)
- Czechia > South Moravian Region
- Brno (0.04)
- Spain > Catalonia
- Asia > Middle East
- UAE (0.14)
- Africa
- Sudan (0.04)
- North Africa (0.04)
- Middle East > Egypt (0.04)
- North America > Canada
- Genre:
- Research Report (0.64)
- Technology: