MARATTO

article

Direct English-to-Yoruba speech Translation model using Transformer with Augmented Attention Mechanism

Abstract

This study explores English-to-Yoruba speech-to-speech translation using a transformer neural network with 36 encoder-decoder pairs and 42 multi-head attention mechanisms. Unlike traditional cascaded models, this approach aligns speech directly, reducing dependency on large spectrogram datasets. Trained on a curated dataset of 200,000+ data points, the model achieved a BLEU score of 0.7, an MOS score of 2.8, and 98% translation accuracy, demonstrating its potential for practical speech translation applications.

Research topics

  • Speech Recognition and Synthesis
  • Speech and Audio Processing

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/nigercon62786.2024.10927001

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.