SERAB: A multi-lingual benchmark for speech emotion recognition

Scheidwasser-Clow, Neil, Kegler, Mikolaj, Beckmann, Pierre, Cernak, Milos

arXiv.org Artificial Intelligence 

Comparing and benchmarking [10] is commonly used for self-supervised pre-training, as different DNN models can often be tedious due to the use of well as a benchmarking method for audio event classification [6, 7, different datasets and evaluation protocols. To facilitate the process, 11]. A recently proposed HEAR challenge [12] focuses on evaluating here, we present the Speech Emotion Recognition Adaptation general-purpose audio representations and extends the concept Benchmark (SERAB), a framework for evaluating the performance underlying AudioSet by including additional tasks. In speech representation and generalization capacity of different approaches for utterancelevel learning, NOSS [6] was recently proposed as a platform SER. The benchmark is composed of nine datasets for SER for evaluating speech-specific feature extractors. It includes diverse in six languages. Since the datasets have different sizes and numbers non-semantic speech processing problems, such as speaker and language of emotional classes, the proposed setup is particularly suitable identification, as well as two SER tasks (CREMA-D [13] and for estimating the generalization capacity of pre-trained DNN-based SAVEE [14]).