MARATTO

article

Arab Twitter Reactions to the Russo-Ukrainian War with Semi-Supervised Learning Techniques

Abstract

The scarcity of labeled datasets poses significant challenges in Natural Language Processing (NLP), particularly for sentiment and partiality analysis of Arabic tweets regarding the Russo-Ukrainian War. Semi-supervised learning (SSL) addresses this issue by leveraging both labeled and unlabeled data. This paper investigates SSL approaches in Arabic NLP by analyzing sentiment and partiality in collected tweets. We trained four models, each with a different approach: (1) a supervised model, (2) a model initialized with pretraining, (3) a model trained with consistency regularization, and (4) a model combining pretraining and consistency regularization. We limited our dataset to 200,000 labeled to compare the effectiveness of the chosen SSL methods with previous work done on a 500,000 tweet dataset. Then we built further over the best model obtained from the previous ones by training it model on 1.3 million labeled tweet, to acheive a state-of-the-art performance 98.9% test accuracy. Our findings demonstrate that SSL can reduce the dependence on labeled data while still achieving competitive performance. This study underscores SSL's potential to enhance Arabic NLP applications, providing more efficient and scalable solutions.

Research topics

  • Public Relations and Crisis Communication

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/niles63360.2024.10753204

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.