conference paper
Today, the amount of electronic text is increasing rapidly in all languages. Since it is common to store several million web pages and hundreds of thousands of text files in electronic devices, manual processing is no longer feasible due to the massive explosion of data. This study presents an automatic keyword extraction approach for the Wolaita language, a low-resource language with rich morphological complexity and phonological distinctiveness. We collected a corpus of 97,745 sentences and applied extensive preprocessing techniques, including text cleaning and stop-word removal, to ensure data quality. We then employed FastText word embeddings for EmbedRank and compared it with TextRank and YAKE! (Yet Another Keyword Extractor), then merged two approaches to enhance keyword extraction. Our experimental results indicate FastText-based methods with a merged approach of EmbedRank and YAKE! achieve higher precision, recall, and F1-score. Specifically, the merged FastText approach demonstrated the best performance with an F1-score of 0.96, highlighting the effectiveness of subword-aware embeddings in handling Wolaita linguistic complexity. This research contributes to low-resource NLP (natural language processing) by providing insights into embedding selection for keyword extraction and establishing a foundation for future enhancements.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/ict4da67218.2025.11282891
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.