dataset · Zenodo (CERN European Organization for Nuclear Research)
A newly expanded evaluation dataset offers human-validated resources for assessing language technologies across Modern Standard Arabic and Moroccan Darija. Spanning 28 validated tables and 11,284 rows, the release includes experimental materials and anonymised annotations across multiple evaluation tracks. The collection covers areas such as lexical semantics, etymological evidence, semantic pairs, comprehension tasks, cross-variety Arabic counterparts, and register continua. It also addresses script variation, code-switching, and digital-service requests. Rather than collapsing linguistic uncertainty into single outputs, the public reference layer explicitly documents multi-reference answers, ambiguity, and stress conditions alongside primary gold standards. This phase-two release supplements an earlier foundational package to provide structured benchmarks for dialectal Arabic processing.
Moroccan Darija differs considerably from Modern Standard Arabic and frequently incorporates mixed phrasing, non-standard scripts, and code-switching. High-quality, human-validated benchmarks are critical for creating digital tools that comprehend everyday dialects. By preserving ambiguity and varied valid interpretations, this resource helps developers measure how reliably language systems handle authentic, complex communication in North African contexts.
Developers building natural language processing tools, automated customer support assistants, and digital public services for North African users can use this dataset to benchmark comprehension and translation accuracy. Because the asset is an evaluation dataset rather than a deployable software product, it serves as an applied testing resource suited for mid-stage model development and validation.
AI-generated from the published abstract. Always read the original work before citing.
Human-validated Phase-2 expansion of the Fusha–Darija evaluation project. The release contains publishable experiment materials and anonymized human annotations for E01, E02, E03, E04, E05, E08, and E09, covering lexical semantics and etymological evidence, Modern Standard Arabic–Moroccan Darija semantic pairs and comprehension tasks, cross-variety Arabic counterparts, framing, register continua, script and code-switching variation, and digital-service requests. The public reference layer preserves primary-gold, multi-reference, ambiguous, stress-condition, and excluded-from-primary statuses rather than collapsing uncertainty. The release contains 28 validated data/annotation/reference tables with 11,284 rows. Raw reviewer workbooks, reviewer identities, private communications, model outputs, and publication results are excluded. This dataset supplements and does not replace the Harvard Dataverse V1 package: https://doi.org/10.7910/DVN/W0C5DV. The archive is mixed-license: project-original materials and derived anonymized annotations are CC BY 4.0; selected third-party components retain CC BY-NC 4.0 or CC BY-SA 4.0 as documented at component and row level within the archive.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.5281/zenodo.21580966
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.