review · American Journal of Artificial Intelligence
Emotion recognition is a core component of affective computing, enabling intelligent systems to interpret human emotional states across critical applications such as healthcare, online education, and human–computer interaction. Early unimodal approaches relying solely on facial expressions, speech, or text have proven insufficient due to noise, cultural variability, and signal ambiguity, prompting a decisive shift toward multimodal integration. This systematic review, conducted following the PRISMA framework, examines this transition by analyzing 89 peer-reviewed studies selected from an initial pool of 160. The objective is to synthesize current methodologies, compare performance across modalities, and identify persistent technical and ethical barriers. Our findings reveal that multimodal systems, which fuse visual, acoustic, linguistic, and physiological signals, consistently outperform unimodal counterparts, achieving accuracy levels above 85% on benchmark datasets. Deep learning architectures particularly convolutional networks for spatial features, recurrent networks for temporal dependencies, and transformer-based models enhanced with attention mechanisms dominate the field, enabling effective dynamic weighting and fusion of heterogeneous data streams. Despite these advances, several challenges impede real-world deployment. Cross-subject and cross-session variability degrades generalizability, while data scarcity and the lack of large-scale, annotated multimodal corpora constrain model training. Computational complexity, especially in transformer-based fusion, limits edge-device feasibility, and ethical concerns surrounding privacy, demographic bias, and model interpretability remain unresolved. Future research must prioritize scalable and lightweight architectures, inclusive and culturally diverse dataset curation, and explainable AI frameworks that build user trust. Ultimately, transitioning these systems from laboratory prototypes to ethically sound, practical applications will require close interdisciplinary collaboration among computer scientists, psychologists, and ethicists, ensuring that emotion recognition technologies are not only accurate but also fair, transparent, and accessible across diverse real-world settings.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.11648/j.ajai.20261002.13
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.