Bridging Modalities: Knowledge Distillation and Masked Training for Translating Multi-Modal Emotion Recognition to Uni-Modal, Speech-Only Emotion Recognition
Muaz, Muhammad, Paull, Nathan, Malagavalli, Jahnavi
–arXiv.org Artificial Intelligence
This paper presents an innovative approach to address the challenges of translating multi-modal emotion recognition models to a more practical and resource-efficient uni-modal counterpart, specifically focusing on speech-only emotion recognition. Recognizing emotions from speech signals is a critical task with applications in human-computer interaction, affective computing, and mental health assessment. However, existing state-of-the-art models often rely on multi-modal inputs, incorporating information from multiple sources such as facial expressions and gestures, which may not be readily available or feasible in real-world scenarios. To tackle this issue, we propose a novel framework that leverages knowledge distillation and masked training techniques.
arXiv.org Artificial Intelligence
Jan-4-2024
- Country:
- North America > United States
- Washington > King County
- Seattle (0.04)
- Texas > Travis County
- Austin (0.05)
- New York > New York County
- New York City (0.04)
- Washington > King County
- Europe > Germany
- Bavaria > Upper Bavaria > Munich (0.04)
- North America > United States
- Genre:
- Research Report
- Promising Solution (0.68)
- New Finding (0.46)
- Research Report
- Industry:
- Education (0.50)
- Health & Medicine > Therapeutic Area
- Psychiatry/Psychology (0.34)
- Technology: