A Cross-Corpus Speech Emotion Recognition Method Based on Supervised Contrastive Learning

minjie, Xiang

arXiv.org Artificial Intelligence 

Recognition of emotion in speech is a key technology in human-computer interaction. With the increasing application of dialogue systems, the demand for SER tasks is also increasing [1]. General deep learning models often require a large amount of data to achieve good results and strong generalization ability, while in the field of SER, there is a lack of large-scale public datasets, and SER tasks also face the problem of weak generalization ability due to the gap between different languages and speakers [2]. The emergence of self-supervised learning speech representation models, which are often trained on hundreds of hours of speech recognition datasets and can be used for numerous downstream tasks related to speech, provides a viable solution to these problems. Many studies [3][4][5] design fine-tuning algorithms for SER tasks based on pre-trained speech representation models.