article
In this paper, we present an extension to the transformer text-to-speech synthesis model that learns prosodic features extracted from a reference expressive speech to synthesize text with the desired prosody. We propose to condition the transformer TTS baseline model with an expressivity encoder, a speaker classifier, and a multi-head cross-attention module used for fusion and alignment between text and expressivity content to achieve the transfer of expressivity from a reference speech to the generated voice. We compare our model to the baseline models of transfer of expressivity namely Global Style Token (GST), Variational AutoEncoder (VAE), and Fine-Grained style control in transformer TTS (FGT). The proposed model shows good performance of expressivity transfer and maintains the quality of the synthesized speech as demonstrated by correlated subjective and objective metrics, listening tests used for evaluating prosody transfer tasks, and naturalness of speech.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/atsip62566.2024.10638975
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.