MARATTO

conference paper

Word Sense Disambiguation Model for Wolaita Language Using Semi-Supervised Learning, Transformer Models, and Explainable AI

Abstract

Human-to-machine interaction has been enhanced by applications, including speech recognition, machine translation, information retrieval, and many natural language processing (NLP) tasks. However, ambiguity remained a challenging problem for it. Word Sense Disambiguation (WSD) is one of the NLP tasks that addresses it by determining the correct sense of a polysemous word based on the context of the text. The Wolaita language consists of several polysemous words. This can pose challenges for high-level NLP applications. This study addressed this issue by developing WSD models. We collected 5,450 sense examples from the Holy Bible, academic texts, online repositories, and news agencies to enhance the generalizability of our model, which learns from diverse language sources. We applied TF-IDF, Word2Vec, and FastText to capture language patterns, and then we trained four semi-supervised learning (SSL) models: self-learning (KNN), co-training (KNN-SVM and KNN-Bagging), and graph-based label propagation. Among these, KNN-SVM with FastText achieved 97.27% accuracy and outperformed other models. Despite the lack of sufficient data, transformer-based models mBERT (88.88% accuracy and 0.1616 loss) and XLM-RoBERTa (88.56% accuracy and 0.1446 loss) achieved encouraging results. We integrated LIME and SHAP explainable AI into our model prediction to uncover influential features that reveal the meaning of ambiguous words.

Research topics

  • Topic Modeling
  • Natural Language Processing Techniques
  • Computational and Text Analysis Methods

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/ict4da67218.2025.11282542

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.