MARATTO

dataset · Zenodo (CERN European Organization for Nuclear Research)

SARCSenti: A Tone-Marked Yoruba Dataset for Sarcasm Detection and Sentiment Classification

2026Open accessLead City University

In plain language

SARCSenti is a newly developed, tone-marked Yoruba dataset designed for affective natural language processing, specifically for binary sarcasm detection and ternary sentiment classification. It comprises 1,507 Yoruba headline records, each annotated for both sarcasm (non-sarcastic or sarcastic) and sentiment (negative, neutral, or positive). The dataset includes 1,334 records derived from BBC Yoruba and 173 Yoruba translations from another news headlines dataset. This resource aims to support research in Yoruba and African low-resource natural language processing, sarcasm detection, sentiment analysis, affective computing, cross-task learning, and the processing of tonal languages. Access to the full dataset is currently restricted due to copyright considerations for some source material.

Key takeaways

  • SARCSenti is a new tone-marked Yoruba dataset for natural language processing.
  • It contains 1,507 Yoruba headline records annotated for both sarcasm and sentiment.
  • The dataset supports binary sarcasm detection and ternary sentiment classification.
  • It is intended to advance research in low-resource African language NLP and affective computing.
  • Access to the full dataset is currently restricted due to copyright considerations for some content.

Why it matters

Understanding the nuances of human language, including sarcasm and sentiment, is vital for effective digital communication. This dataset provides a crucial resource for developing artificial intelligence that can better interpret the Yoruba language, especially its tonal characteristics, thereby improving technology's ability to process and respond to African languages more accurately.

Commercialisation angle

This dataset is an early-stage research tool, providing foundational data for developing advanced natural language processing applications for Yoruba. It could enable future tools for sentiment analysis, content moderation, or intelligent chatbots in Yoruba. Potential users include researchers and developers working on African language technologies, as well as organisations needing to analyse public opinion or detect specific tones in Yoruba text. Full commercialisation pathways are not indicated, as the dataset itself is a research enabler.

AI-generated from the published abstract. Always read the original work before citing.

Abstract

SARCSenti is a tone-marked Yoruba dataset developed for affective natural language processing, specifically binary sarcasm detection and ternary sentiment classification. The analytical release contains 1,507 Yoruba headline records, annotated for both sarcasm and sentiment. Sarcasm labels are non-sarcastic (0) and sarcastic (1), while sentiment labels are negative (0), neutral (1), and positive (2). The dataset contains 1,334 BBC Yoruba-derived records and 173 Yoruba translations derived from the Misra/Kaggle News Headlines Dataset for Sarcasm Detection. Record-level provenance is provided in the source field. Three records from the original research workbook were excluded from the analytical release because of missing or invalid numeric sentiment labels. SARCSenti was developed as part of research conducted in the Department of Computer Science, Lead City University, Ibadan, Nigeria. The dataset is intended to support research on Yoruba and African low-resource NLP, sarcasm detection, sentiment analysis, affective computing, cross-task learning, and the processing of tonal languages. Yoruba text is provided in UTF-8 and normalized using Unicode NFC. The accompanying README and data dictionary provide information on dataset structure, labels, provenance, and usage. Access to the full-text dataset is restricted because a substantial portion contains source-derived BBC Yoruba headline text for which open redistribution rights have not yet been confirmed. Research access may be considered subject to applicable source terms and copyright restrictions.

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.5281/zenodo.22308685

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.