Exploiting Dialect Identification in Automatic Dialectal Text Normalization
Alhafni, Bashar, Al-Towaity, Sarah, Fawzy, Ziyad, Nassar, Fatema, Eryani, Fadhl, Bouamor, Houda, Habash, Nizar
–arXiv.org Artificial Intelligence
Dialectal Arabic is the primary spoken language used by native Arabic speakers in daily communication. The rise of social media platforms has notably expanded its use as a written language. However, Arabic dialects do not have standard orthographies. This, combined with the inherent noise in user-generated content on social media, presents a major challenge to NLP applications dealing with Dialectal Arabic. In this paper, we explore and report on the task of CODAfication, which aims to normalize Dialectal Arabic into the Conventional Orthography for Dialectal Arabic (CODA). We work with a unique parallel corpus of multiple Arabic dialects focusing on five major city dialects. We benchmark newly developed pretrained sequence-to-sequence models on the task of CODAfication. We further show that using dialect identification information improves the performance across all dialects. We make our code, data, and pretrained models publicly available.
arXiv.org Artificial Intelligence
Jul-3-2024
- Country:
- Africa > Middle East
- Egypt > Cairo Governorate
- Cairo (0.05)
- Morocco (0.04)
- Tunisia > Tunis Governorate
- Tunis (0.05)
- Egypt > Cairo Governorate
- Asia
- China
- India > Maharashtra
- Mumbai (0.04)
- Japan
- Honshū
- Chūbu > Aichi Prefecture
- Nagoya (0.04)
- Kansai > Osaka Prefecture
- Osaka (0.04)
- Chūbu > Aichi Prefecture
- Kyūshū & Okinawa > Kyūshū
- Miyazaki Prefecture > Miyazaki (0.05)
- Honshū
- Middle East
- Lebanon > Beirut Governorate
- Beirut (0.05)
- Qatar > Ad-Dawhah
- Doha (0.05)
- Republic of Türkiye > Istanbul Province
- Istanbul (0.04)
- UAE > Abu Dhabi Emirate
- Abu Dhabi (0.14)
- Lebanon > Beirut Governorate
- Singapore (0.05)
- Europe
- Belgium > Brussels-Capital Region
- Brussels (0.04)
- United Kingdom
- England > Oxfordshire
- Oxford (0.04)
- Scotland > City of Edinburgh
- Edinburgh (0.04)
- England > Oxfordshire
- Ireland > Leinster
- County Dublin > Dublin (0.04)
- Croatia > Dubrovnik-Neretva County
- Dubrovnik (0.04)
- Middle East > Republic of Türkiye
- Istanbul Province > Istanbul (0.04)
- Netherlands > South Holland
- Dordrecht (0.04)
- Ukraine > Kyiv Oblast
- Kyiv (0.04)
- Spain > Catalonia
- Barcelona Province > Barcelona (0.04)
- Slovenia (0.04)
- Denmark > Capital Region
- Copenhagen (0.04)
- France > Provence-Alpes-Côte d'Azur
- Bouches-du-Rhône > Marseille (0.05)
- Italy > Tuscany
- Florence (0.04)
- Iceland > Capital Region
- Reykjavik (0.05)
- Belgium > Brussels-Capital Region
- North America
- Canada > Quebec
- Montreal (0.04)
- United States
- California (0.04)
- Louisiana > Orleans Parish
- New Orleans (0.04)
- Maryland (0.04)
- Michigan > Washtenaw County
- Ann Arbor (0.04)
- New Mexico > Santa Fe County
- Santa Fe (0.04)
- New York (0.04)
- Pennsylvania > Allegheny County
- Pittsburgh (0.04)
- Washington > King County
- Seattle (0.14)
- Canada > Quebec
- Africa > Middle East
- Genre:
- Research Report > New Finding (0.46)
- Technology: