Goto

Collaborating Authors

 perspective loss




Towards Diverse Perspective Learning with Selection over Multiple Temporal Poolings

arXiv.org Artificial Intelligence

In Time Series Classification (TSC), temporal pooling methods that consider sequential information have been proposed. However, we found that each temporal pooling has a distinct mechanism, and can perform better or worse depending on time series data. We term this fixed pooling mechanism a single perspective of temporal poolings. In this paper, we propose a novel temporal pooling method with diverse perspective learning: Selection over Multiple Temporal Poolings (SoM-TP). SoM-TP dynamically selects the optimal temporal pooling among multiple methods for each data by attention. The dynamic pooling selection is motivated by the ensemble concept of Multiple Choice Learning (MCL), which selects the best among multiple outputs. The pooling selection by SoM-TP's attention enables a non-iterative pooling ensemble within a single classifier. Additionally, we define a perspective loss and Diverse Perspective Learning Network (DPLN). The loss works as a regularizer to reflect all the pooling perspectives from DPLN. Our perspective analysis using Layer-wise Relevance Propagation (LRP) reveals the limitation of a single perspective and ultimately demonstrates diverse perspective learning of SoM-TP. We also show that SoM-TP outperforms CNN models based on other temporal poolings and state-of-the-art models in TSC with extensive UCR/UEA repositories.


An Empirical Study and Improvement for Speech Emotion Recognition

arXiv.org Artificial Intelligence

The above studies all focus on the emotional recognition of a single utterance. Considering the relevance between Multimodal speech emotion recognition aims to detect speakers' utterances, more recent works step towards the actual scenario, emotions from audio and text. Prior works mainly focus conversational multimodal emotion recognition. Poria on exploiting advanced networks to model and fuse different et al. [10] adopt hierarchical modeling, and they extract modality information to facilitate performance, while neglecting utterance-level features first and then aggregate dialoguelevel the effect of different fusion strategies on emotion information. In the following work [1], they propose an recognition. In this work, we consider a simple yet important attention-based fusion (AT-Fusion) approach [1] to obtain different problem: how to fuse audio and text modality information is modal utterance representations and get the global information more helpful for this multimodal task. Further, we propose by additional LSTM [11] and self-attention mechanism a multimodal emotion recognition model improved by perspective [12]. Besides, graph neural networks are also employed loss. Empirical results show our method obtained to fuse audio and text modality features [13, 14].