MARATTO

article · Scientific Reports

TigCLaF: a cross-lingual large language model framework for sentiment-aware text classification in low-resource tigrigna

2026Open accessMekelle University

Abstract

This paper introduces TigCLaF, a novel cross-lingual large language model framework for sentiment-aware text classification in low-resource Tigrigna. The framework integrates tokenizer extension with adaptation, continual pretraining on unlabeled Tigrigna data, and LoRA-based parameter-efficient fine-tuning to enable effective cross-lingual adaptation from high-resource languages. Leveraging recent advances in multilingual pre-trained language models and large language models, we investigate zero-shot, few-shot, and full fine-tuning strategies for sentiment detection, incorporating transformer models XLM-RoBERTa and AfriBERTa, as well as instruction-tuned LLaMA models. Our approach integrates a Tigrigna sentiment lexicon into transformer-based embeddings via feature fusion, thereby enhancing preservation of the sentiment signal during cross-lingual transfer. The proposed framework is evaluated on a newly curated dataset of 30,000 Tigrigna instances, supported by auxiliary English and Amharic sentiment corpora for transfer learning. Experimental results show that sentiment-aware feature integration improves classification accuracy and Macro-F1 up to 7% over baseline multilingual models without sentiment augmentation. Furthermore, parameter-efficient fine-tuning LoRA achieved competitive accuracy while reducing model size and inference latency, making it suitable for computationally constrained settings. The systematic error analysis highlights the roles of script-specific preprocessing, idiomatic expressions, and nuances of cultural sentiment in classification performance. The proposed framework demonstrates the viability of combining LLM-based cross-lingual transfer with masked sentiment-aware enhancements for practical, resource-efficient NLP applications in low-resource language contexts.

Research topics

  • Sentiment Analysis and Opinion Mining
  • Topic Modeling
  • Hate Speech and Cyberbullying Detection

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1038/s41598-026-42786-4

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.