Advancements in Arabic Grammatical Error Detection and Correction: An Empirical Investigation
Alhafni, Bashar, Inoue, Go, Khairallah, Christian, Habash, Nizar
–arXiv.org Artificial Intelligence
Grammatical error correction (GEC) is a well-explored problem in English with many existing models and datasets. However, research on GEC in morphologically rich languages has been limited due to challenges such as data scarcity and language complexity. In this paper, we present the first results on Arabic GEC using two newly developed Transformer-based pretrained sequence-to-sequence models. We also define the task of multi-class Arabic grammatical error detection (GED) and present the first results on multi-class Arabic GED. We show that using GED information as an auxiliary input in GEC models improves GEC performance across three datasets spanning different genres. Moreover, we also investigate the use of contextual morphological preprocessing in aiding GEC systems. Our models achieve SOTA results on two Arabic GEC shared task datasets and establish a strong benchmark on a recently created dataset. We make our code, data, and pretrained models publicly available.
arXiv.org Artificial Intelligence
Nov-9-2023
- Country:
- Africa > Middle East
- Morocco (0.04)
- Asia
- China
- Japan > Kyūshū & Okinawa
- Kyūshū > Miyazaki Prefecture > Miyazaki (0.04)
- Middle East
- Qatar > Ad-Dawhah
- Doha (0.04)
- Republic of Türkiye > Istanbul Province
- Istanbul (0.04)
- UAE > Abu Dhabi Emirate
- Abu Dhabi (0.04)
- Qatar > Ad-Dawhah
- Taiwan > Taiwan Province
- Taipei (0.04)
- Europe
- Switzerland > Geneva
- Geneva (0.04)
- Czechia > Prague (0.04)
- Belgium > Brussels-Capital Region
- Brussels (0.04)
- Ireland > Leinster
- County Dublin > Dublin (0.04)
- Middle East > Republic of Türkiye
- Istanbul Province > Istanbul (0.04)
- Ukraine > Kyiv Oblast
- Kyiv (0.04)
- United Kingdom > England
- Cambridgeshire > Cambridge (0.04)
- Greater Manchester > Manchester (0.04)
- Bulgaria > Sofia City Province
- Sofia (0.04)
- Netherlands (0.04)
- France > Provence-Alpes-Côte d'Azur
- Bouches-du-Rhône > Marseille (0.04)
- Italy > Tuscany
- Florence (0.04)
- Germany > Berlin (0.04)
- Iceland > Capital Region
- Reykjavik (0.04)
- Switzerland > Geneva
- North America
- Canada
- British Columbia > Metro Vancouver Regional District
- Vancouver (0.04)
- Quebec > Montreal (0.04)
- British Columbia > Metro Vancouver Regional District
- Dominican Republic (0.04)
- United States
- Washington > King County
- Seattle (0.14)
- Georgia > Fulton County
- Atlanta (0.04)
- Oregon > Multnomah County
- Portland (0.04)
- Louisiana > Orleans Parish
- New Orleans (0.04)
- New York (0.04)
- Maryland > Baltimore (0.04)
- Colorado > Denver County
- Denver (0.04)
- California > San Diego County
- San Diego (0.04)
- Minnesota > Hennepin County
- Minneapolis (0.14)
- Texas > Travis County
- Austin (0.04)
- Washington > King County
- Canada
- Africa > Middle East
- Genre:
- Research Report > New Finding (0.46)
- Technology: