article
In this work, we build upon our previous research where we introduced phonological features as input to text-to-speech systems. While the use of phonological features is not a novel concept in our research, our focus in this study is on the comprehensive analysis of the embeddings produced by the encoder model, which we believe offers novel insights into the model's ability to capture and generalize phonological patterns across languages. Cross-lingual transfer experiments are conducted using both a resource-rich and a resource-constrained language to explore the model's cross-lingual transfer capabilities across different linguistic families. The analysis of the embedding vectors produced by the encoder model is conducted using cluster maps to visualize the hierarchical clusters obtained using a clustering procedure. This analysis reveals the model's learning patterns and provides insights into how phonological features contribute to the model's ability to handle linguistic diversity and data scarcity.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/tsp63128.2024.10605769
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.