MARATTO

dataset · Mendeley Data

RIYE Textual and Geospatial Corpus: An Aggregated Multidialectal Dataset for Low-Resource Language Processing

Abstract

This dataset integrates textual and spatial metadata collected under the Digiculture RIYE Project framework to support research in computational linguistics, spatial humanities, and low-resource dialect analysis. The corpus comprises three complementary datasets: RIYE_CSC-HTM_Aggregate.csv, which aggregates raw, qualitative text entries across cultural categories (such as food, festivals, music, history, and folklore) collected by the Computer Science (CSC) and Hospitality and Tourism Management (HTM) groups from the Federal University of Agriculture, Abeokuta; RIYE_CSC_Dataset.csv, which provides a structured matrix of thematic cultural elements (including dialect, clothing, religion, leadership, and performance arts) mapped across Local Government Areas (LGAs), towns/wards, and GIS location coordinates; and ogun_state_lgas_wards_coordinates.csv containing constitutionally recognised LGAs, wards, and coordinates extracted from the publicly available database compiled by Independent National Electoral Commission (INEC) before the 2023 General Elections. By combining unstructured descriptive text with spatially anchored domain attributes across regional zones, this collection establishes a rich baseline for requirement analysis, computational text processing, and geospatial modelling of dialectal and cultural elements.

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.17632/v9gdz3k7zg

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.