article
The availability of large annotated corpora remains a major challenge for the development of natural language processing systems for underresourced languages such as Arabic.In this paper, we present two annotated corpora dedicated to Modern Standard Arabic.These corpora are open-source and freely available on the Hugging Face platform.The first corpus, annotated by theme and designed to provide a balanced representation of contemporary Arabic usage, comprises approximately 76 million words collected from diverse sources covering multiple domains and geographical regions.The second corpus, containing approximately one million words, is a sub-corpus extracted from the first.It was annotated with lemma tags using a semi-automatic approach that combines automatic annotation with the Alkhalil lemmatizer and MADAMIRA, followed by manual validation.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.18653/v1/2026.abjadnlp-1.27
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.