MARATTO

article

Leveraging Transformer-Based Models for Cyberbullying Detection in the Moroccan Arabic Dialect

Abstract

Cyberbullying is an increasing threat on social media, with serious consequences for mental health, particularly among young people. Despite growing global efforts to address this problem, developing accurate detection systems remains challenging for low-resource languages in both text and audio modalities, such as Arabic and its dialects. This paper presents a binary-labeled dataset for cyberbullying detection in Moroccan Darija. The dataset merges over 4,000 newly collected YouTube comments with the OMCD (Offensive Moroccan Comments Dataset) corpus of 8,024 comments, originally labeled for offensive language. To better reflect the nature of cyberbullying, the entire corpus was reannotated from scratch, clearly distinguishing between bullying and non-bullying content, including subtle forms like sarcasm, shaming, and indirect aggression. Several Arabic transformer models were fine-tuned and evaluated using standard classification metrics. The results show that dialectspecific models, particularly DarijaBERT, achieve the best performance, underlining the importance of context-aware annotation and in-domain pre-training for cyberbullying detection in lowresource settings.

Research topics

  • Hate Speech and Cyberbullying Detection
  • Bullying, Victimization, and Aggression
  • Authorship Attribution and Profiling

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/wincom65874.2025.11313421

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.