An Empirical Study and Improvement for Speech Emotion Recognition
Wu, Zhen, Lu, Yizhe, Dai, Xinyu
–arXiv.org Artificial Intelligence
The above studies all focus on the emotional recognition of a single utterance. Considering the relevance between Multimodal speech emotion recognition aims to detect speakers' utterances, more recent works step towards the actual scenario, emotions from audio and text. Prior works mainly focus conversational multimodal emotion recognition. Poria on exploiting advanced networks to model and fuse different et al. [10] adopt hierarchical modeling, and they extract modality information to facilitate performance, while neglecting utterance-level features first and then aggregate dialoguelevel the effect of different fusion strategies on emotion information. In the following work [1], they propose an recognition. In this work, we consider a simple yet important attention-based fusion (AT-Fusion) approach [1] to obtain different problem: how to fuse audio and text modality information is modal utterance representations and get the global information more helpful for this multimodal task. Further, we propose by additional LSTM [11] and self-attention mechanism a multimodal emotion recognition model improved by perspective [12]. Besides, graph neural networks are also employed loss. Empirical results show our method obtained to fuse audio and text modality features [13, 14].
arXiv.org Artificial Intelligence
Apr-7-2023
- Country:
- Asia > China > Jiangsu Province > Nanjing (0.05)
- Genre:
- Research Report > New Finding (0.34)
- Technology: