Enhanced Human Emotion Recognition through Multimodal Data using Deep Learning and Late Fusion Technique
نویسندگان
1 Department of Information Technology, Pune Institute of Computer Technology, Pune, Maharashtra, India
2 Department of Computer Science and Engineering, MIT SOC, MIT Art, Design and Technology University, Pune, India
doi
10.5829/ije.2026.39.09c.09چکیده
Emotion recognition is pivotal for advancing artificial intelligence (AI) and human-computer interaction (HCI), enabling systems to better understand and respond to human needs. Traditional approaches often rely on single modalities, such as speech or facial expressions, which are limited by environmental noise, occlusions, and viewpoint variations. To address these challenges, this study proposes a robust multimodal deep learning framework that integrates Long Short-Term Memory (LSTM) networks for sequential audio analysis with ResNet50 for facial emotion recognition. A late fusion strategy is employed to merge modality-specific outputs, leveraging the complementary strengths of both modalities. The proposed method is evaluated on benchmark multimodal datasets, RAVDESS and MELD, using metrics such as classification accuracy and emotion-specific performance. Comparative analysis is performed against both unimodal baselines (LSTM for audio, ResNet101 for facial images) and state-of-the-art multimodal fusion techniques, including early fusion, hybrid fusion, and weighted decision-level fusion methods. Experimental results show that the proposed ResNet50+LSTM model with late fusion achieves 90.03% accuracy on RAVDESS and 88.57% on MELD, outperforming existing approaches across most emotion categories, particularly “Happy” and “Anger,” which demonstrate higher precision and recall compared to “Sadness” and “Surprise,” with slightly lower F1 scores. These findings validate the framework’s robustness and its potential for real-world applications in emotion-aware AI systems across domains such as mental health, education, customer service, and adaptive interfaces.