Computer Science > Computation and Language

arXiv:2305.14734 (cs)

[Submitted on 24 May 2023 (v1), last revised 9 Nov 2023 (this version, v2)]

Title:Advancements in Arabic Grammatical Error Detection and Correction: An Empirical Investigation

Authors:Bashar Alhafni, Go Inoue, Christian Khairallah, Nizar Habash

View PDF

Abstract:Grammatical error correction (GEC) is a well-explored problem in English with many existing models and datasets. However, research on GEC in morphologically rich languages has been limited due to challenges such as data scarcity and language complexity. In this paper, we present the first results on Arabic GEC using two newly developed Transformer-based pretrained sequence-to-sequence models. We also define the task of multi-class Arabic grammatical error detection (GED) and present the first results on multi-class Arabic GED. We show that using GED information as an auxiliary input in GEC models improves GEC performance across three datasets spanning different genres. Moreover, we also investigate the use of contextual morphological preprocessing in aiding GEC systems. Our models achieve SOTA results on two Arabic GEC shared task datasets and establish a strong benchmark on a recently created dataset. We make our code, data, and pretrained models publicly available.

Comments:	Accepted to EMNLP 2023
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2305.14734 [cs.CL]
	(or arXiv:2305.14734v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2305.14734

Submission history

From: Bashar Alhafni [view email]
[v1] Wed, 24 May 2023 05:12:58 UTC (7,481 KB)
[v2] Thu, 9 Nov 2023 16:10:59 UTC (7,540 KB)

Computer Science > Computation and Language

Title:Advancements in Arabic Grammatical Error Detection and Correction: An Empirical Investigation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Advancements in Arabic Grammatical Error Detection and Correction: An Empirical Investigation

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators