MARATTO

preprint · Zenodo (CERN European Organization for Nuclear Research)

WDYW.01: Tone Restoration for Standard Yorùbá

Abstract

Purpose: This paper addresses automatic tone restoration for Standard Yorùbá, a three-tone language whose digital texts are predominantly written without tone marks. It presents ToneMarker, the first module (WDYW.01) of the Well-formed Digital Yorùbá Writing (WDYW) research programme, and evaluates its performance on an independent test set. Design/methodology/approach: ToneMarker is a freely accessible, browser-based tool built on a verified corpus of 11.8 million words. It employs a four-layer architecture combining expert-verified corrections, sentence-level matching, context-based statistical modelling, and frequency-based prediction, augmented by a grammar-based disambiguation layer for selected high-frequency ambiguities. Performance was evaluated on a gold-standard test set of 48 Yorùbá Wikipedia sentences independently tone-marked by the author. Findings: ToneMarker achieved 77.1% word-level accuracy (95% CI: 74.1–79.8%), increasing to 78.3% when the grammatically irreducible ni/ní homophony was excluded. The study also identifies and resolves a previously under-documented double-diacritic-stripping bug that silently corrupts Yorùbá text in standard tokenisation pipelines and demonstrates that Unicode normalisation is essential for accurate evaluation of tone-restoration systems. Originality: Within the author’s knowledge, this is the first Yorùbá tone-restoration system to combine corpus-scale statistical modelling with a grammar-based disambiguation layer. It is also, as far as we are aware, the first study to document and quantify the effects of double-diacritic stripping in Yorùbá NLP pipelines. Contribution to the field: The paper provides a freely deployable tool for Yorùbá scholars, students, and language professionals while establishing a foundation for future lexical, grammatical, and neural approaches to digital Yorùbá writing support.

Research topics

  • Phonetics and Phonology Research
  • Linguistic Variation and Morphology
  • Speech Recognition and Synthesis

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.5281/zenodo.21422029

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.