article
This study addresses the challenge of multi-label text classification in the Arabic language, focusing on movie genre categorization using plot summaries. Even though over 400 million people speak Arabic, its natural language processing (NLP) advances are not keeping up with those of other languages because of data shortages and quality difficulties. Three key contributions are made by this research to narrow this gap: a thorough analysis of prior research on Arabic multi-label text classification; the introduction of a newly curated dataset containing 22 genre labels for Egyptian movies; and the creation of a framework for optimizing large language models (LLMs) on this dataset. Our methodology leverages multiple fine-tuned BERT models to evaluate and compare performance against existing datasets, offering a reproducible path for future research.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/niles63360.2024.10753176
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.