MARATTO

article

Bimodal Emotional Recognition based on Long Term Recurrent Convolutional Network

Abstract

Determining a person’s emotional state remains a non-trivial task relating to the ambiguous definition of the emotion itself and the different tools used to identify emotional aspects. In this study, we propose an Emotional Recognition System by adopting a multimodal approach that combines the based speech emotion features and the based facial expressions features. Accordingly, the proposed recognition system contains three parts. The first component will be reserved to extract the facial expressions features using a deep learning-based network called Long-Term Recurrent Convolutional Network (LRCN). The second component will be set aside to extract speech emotion features using the Alex-Net deep learning-based network. Finally, we dedicated the last component to combine the information learned from the two previous modalities through the late fusion. In order to evaluate the performance of the proposed approach, we tested it with the Ryerson Audio Visual Database of Emotional Speech and Song (RAVDESS) human emotions. Obtained results show that the accuracy rate was improved by using the fusion strategy. Indeed, the accuracy rate increased from 78.82 (for facial modality) and 76.39 (for speech modality) to 85.76.

Research topics

  • Emotion and Mood Recognition
  • Speech and Audio Processing
  • Music and Audio Processing

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1145/3607720.3607740

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.