MARATTO

article · Procedia Computer Science

Unveiling embedded features in Wav2vec2 and HuBERT msodels for Speech Emotion Recognition

202417 citationsOpen accessHassan II University Casablanca

Abstract

Speech Emotion Recognition (SER) is a very interesting task that allows the machine to identify and recognize the different emotional states from human speech using new technologies. The SER can be represented by two main steps, namely feature extraction and emotion classification. Our contribution to the SER field will focus on these two phases. This paper seeks to investigate the influence of embedded features in Wav2vec2 and HuBERT models on SER, two variants per module are implemented, including Wav2vec2 base, Wav2vec2 large, HuBERT large and HuBERT X-large. In addition, we adopt a linear Support Vector Machine (SVM) as a downstream model to recognize emotions. The proposed approach relying on the combination of HuBERT X-large features with the SVM model led to the highest recognition rate of 82.6% on the RAVDESS database. Moreover, the results obtained are promising and in compliance with the current SER state-of-the-art. Hence, the embedded features of HuBERT X-large model have shown significant results for SER.

Research topics

  • Emotion and Mood Recognition
  • Speech and Audio Processing
  • Music and Audio Processing

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1016/j.procs.2024.02.074

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.