MARATTO

article

Hybrid LLM and Rule-Based Synthetic Data Generation for Arabic Grammatical Error Correction

20251 citationAlexandria University

Abstract

The correction of Arabic grammatical errors has re-cently gained interest due to the emergence of strong transformer models trained on internet-scale Arabic data and large language models capable of generalizing well to new tasks and languages. However, Arabic grammatical error correction still suffers from data scarcity, and the available datasets are insufficient to train efficient transformer models capable of performing well on all types of Arabic grammatical errors. In this work, we introduce a scalable recipe to generate synthetic data for Arabic grammatical error correction and control the distribution of error types with high granularity to boost performance in underrepresented error types in the original training data. Models trained on our synthesized data demonstrate enhanced performance in addressing challenging Arabic grammatical errors compared to those trained without it, and they rival state-of-the-art Arabic grammatical error correction systems.

Research topics

  • Natural Language Processing Techniques
  • Educational Technology and Assessment
  • Text Readability and Simplification

Sustainable Development Goals

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/icmisi65108.2025.11115884

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.