article
The rapid advancement of internet communication technologies has enabled users to publish news and share opinions widely. While this brings convenience, it also facilitates the spread of fake news. Although extensive research has been conducted on high-resource languages such as English, Arabic—particularly its dialects—remains underexplored due to its complex morphology, lexical ambiguity, and the scarcity of annotated datasets.We investigate the effectiveness of pretrained language models (PLMs) and large language models (LLMs) for fake news detection in the Algerian dialect using the FASSILA dataset. We evaluate a diverse set of PLMs, including the multilingual transformer (mBERT), Arabic-centric transformers (AraBERT, MARBERT, CAMeLBERT-DA, AraELECTRA), and a dialect-specific encoder (DZiriBERT), as well as LLMs (AraGPT2 variants and LLaMA 3). Our experiments compare multilingual, Arabic-specific, and dialect-adapted models, with a focus on analyzing the impact of dialectal pretraining and model scale.The results show that AraGPT2 variants outperform all other models in our experiments, achieving an improvement of approximately 2–3% compared to previously reported results on the FASSILA dataset. These findings highlight the value of large-scale Arabic pretraining for capturing richer semantic patterns and improving performance in downstream NLP tasks, including dialectal fake news detection.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/rif68108.2025.11406852
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.