article · FUDMA Journal of Sciences
Emotion recognition systems often struggle when relying on a single data source or unbalanced training datasets. To overcome these limitations, a multi-modal model combines facial expression analysis with physiological signal processing. The framework uses Convolutional Neural Networks for facial image analysis and Long Short-Term Memory networks for physiological signals, merging both streams via hybrid fusion. Generative Adversarial Networks augment the FER-2013 facial expression and DEAP physiological datasets, addressing class imbalances and expanding feature diversity. Testing showed that data augmentation significantly boosted performance across individual models. The combined multi-modal architecture reached 93 percent accuracy and a 92 percent F1-score when trained on the augmented data, surpassing single-modality alternatives. This approach demonstrates how synthetic data generation and multi-modal integration can improve the reliability of emotion detection systems across complex emotional states.
Emotion recognition technology underpins advances in healthcare monitoring, security systems, entertainment, and human-computer interfaces. By combining facial cues with internal physiological signals, systems become significantly less prone to error and can better identify subtle or underrepresented emotions. Demonstrating that synthetic data generation via generative models overcomes data scarcity provides a practical path to training more dependable artificial intelligence tools.
Potential applications span healthcare diagnostics, user experience design in human-computer interaction, interactive entertainment, and security monitoring. Relevant users include software developers building affective computing tools and clinical or digital media teams. Because the research was evaluated entirely on standard benchmark datasets, the technology is at an early experimental stage and requires real-world testing, integration with physical sensors, and attention to ethical considerations before commercial deployment.
AI-generated from the published abstract. Always read the original work before citing.
Emotion recognition is a critical area of research with applications in healthcare, human-computer interaction (HCI), security, and entertainment. This study addressed the limitations of single-modal emotion recognition systems by developing a multi-modal emotion recognition model that integrates facial expressions and physiological signals, enhanced by Generative Adversarial Networks (GANs). It aims at improving accuracy, reliability, and robustness in emotion detection, particularly underrepresented emotions. The study utilized the FER-2013 dataset for facial expressions and the DEAP dataset for physiological signals. GANs were employed to augment datasets, address class imbalances and enhance feature diversity. A hybrid multi-modal model was developed, combining Convolutional Neural Networks (CNNs) for facial expression recognition and Long Short-Term Memory (LSTM) networks for physiological signal analysis. Hybrid fusion was used to integrate features at multiple levels, maximizing the complementary strengths of each modality. The results demonstrate significant improvements in emotion recognition. Without GAN augmentation, the CNN and LSTM models achieved accuracies of 62% and 76%, respectively. The hybrid model outperformed, gaining 90% across all metrics. With GAN-augmented datasets, the CNN and LSTM models improved to 81% and 86%, respectively, while the hybrid (multi-modal) model achieved state-of-the-art performance with 93% accuracy and an F1-score of 92%. These findings underscore the efficacy of GANs in enhancing data diversity and the advantages of multi-modal integration for robust emotion recognition. The study contributes to knowledge by introducing a GAN-augmented hybrid multi-modal framework, advancing methodologies in emotion recognition. Recommendations for future work include addressing ethical considerations in emotion recognition systems.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.33003/fjs-2025-0905-3412
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.