MARATTO

article · IEEE Access

RFGETT-TTS: Robust Fine-Grained Expressivity Transfer With Transformer for Text-to-Speech Synthesis

Abstract

Neural text-to-speech (TTS) research has advanced significantly, yielding various approaches that generate speech with enhanced naturalness. Despite these strides, synthesizing expressive speech remains a significant challenge due to the complex and variable nature of explicit human prosody. In this paper, we presented an effective TTS approach that transfers expressivity from a reference speech to a target spoken text. The proposed approach, Robust Fine-Grained Expressivity Transfer with Transformer, RFGETT-TTS, extends and enhances the Transformer-based text-to-speech architecture to enable transfer of expressivity from a reference utterance to synthesized speech. Key components of RFGETT-TTS include: (1) a prosody extraction module that aligns prosodic features with phoneme inputs, providing stable prosodic representations to the NTTS system, (2) an expressivity aggregation module that encodes and fuses expressive characteristics into a single expressive vector using convolutional layers and normalization, (3) a Multi-Head Cross-Attention mechanism that aligns and integrates linguistic inputs with expressivity features derived from the reference utterance, and (4) the speaker module that encodes speaker identity representation. The Transformer encoder's output is then fused with speaker embeddings and used to condition the Transformer decoder. Extensive evaluations on ESD and EmoV-DB datasets demonstrate that RFGETT-TTS outperforms four TTS models in expressivity transfer while maintaining high-quality of synthesized speech, validated by both objective and subjective evaluation metrics. Objective metric like Mel Cepstral Distortion shows that RFGETT-TTS achieves better than four models with a gap up to 0.18, while statistical subjective results on Expressivity Mean Opinion Score show significant difference (p-value < 0.05) than four models.

Research topics

  • Speech Recognition and Synthesis
  • Phonetics and Phonology Research
  • Emotion and Mood Recognition

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/access.2025.3648672

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.