MARATTO

article

Expressivity Transfer In Transformer-Based Text-To-Speech Synthesis

Abstract

In this paper, we present an extension to the transformer text-to-speech synthesis model that learns prosodic features extracted from a reference expressive speech to synthesize text with the desired prosody. We propose to condition the transformer TTS baseline model with an expressivity encoder, a speaker classifier, and a multi-head cross-attention module used for fusion and alignment between text and expressivity content to achieve the transfer of expressivity from a reference speech to the generated voice. We compare our model to the baseline models of transfer of expressivity namely Global Style Token (GST), Variational AutoEncoder (VAE), and Fine-Grained style control in transformer TTS (FGT). The proposed model shows good performance of expressivity transfer and maintains the quality of the synthesized speech as demonstrated by correlated subjective and objective metrics, listening tests used for evaluating prosody transfer tasks, and naturalness of speech.

Research topics

  • Speech Recognition and Synthesis
  • Natural Language Processing Techniques
  • Speech and dialogue systems

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/atsip62566.2024.10638975

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.