Self-Supervised Audio-Visual Representation Learning with Relaxed Cross-Modal Synchronicity

Sarkar, Pritam, Etemad, Ali

arXiv.org Artificial Intelligence 

We refer to this as'asynchronous In recent years, self-supervised learning has shown great cross-modal' optimization, a concept that has not promise in learning strong representations without humanannotated been explored in prior works. We use 3 datasets of different labels (Chen et al. 2020; Chen and He 2021; sizes: Kinetics-Sound (Arandjelovic and Zisserman 2017), Caron et al. 2018), and emerged as a strong competitor for Kinetics400 (Kay et al. 2017), and AudioSet (Gemmeke fully-supervised pretraining. There are a number of benefits et al. 2017), to pretrain CrissCross. We evaluate CrissCross to such methods. Firstly, they reduce the time and resources on different downstream tasks, namely action recognition, required for expensive human annotations and allow researchers sound classification, and action retrieval. We use 2 popular to directly use large uncurated datasets for learning benchmarks UCF101 (Soomro, Zamir, and Shah 2012) and meaningful representations. Moreover, the models trained HMDB51 (Kuehne et al. 2011) to perform action recognition in a self-supervised fashion learn more abstract representations, and retrieval, while ESC50 (Piczak 2015) and DCASE which are useful for a variety of downstream tasks (Stowell et al. 2015) are used for sound classification.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found