MARATTO

article

Unlocking Additional Learning Capabilities of Whisper for Arabic Language Via Instruction Fine-tuning

Abstract

Encoder-decoder models based on transformers showed remarkable results in different natural language processing (NLP) generative tasks, including automatic speech recognition (ASR). In this paper, we focus on one of the most popular encoder-decoder models, Whisper, to increase its capabilities in Arabic language using instruction fine-tuning. We propose an Arabic dataset featuring diverse tasks with different prompts directly applied to acoustic features. The proposed dataset is synthesized from popular Arabic speech datasets covering various tasks such as gender identification, audio environment identification, and audio-text alignment. Moreover, we apply instruction fine-tuning to Whisper on the synthesized dataset, and our results demonstrate that Whisper can acquire new capabilities with minimal or no significant increase in word error rate (WER) and character error rate (CER). Specifically, the large Whisper model achieves an F1-score of 0.98 for gender identification, 0.76 for audio environment identification, and an Intersection over Union (IoU) of 0.86 for audio-text alignment. These enhancements are achieved at the cost of a slight increase in WER by 0.0085 and in CER by 0.0038.

Research topics

  • Speech Recognition and Synthesis
  • Natural Language Processing Techniques
  • Speech and Audio Processing

Sustainable Development Goals

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/aiccsa66935.2025.11315296

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.