About contrastive unsupervised representation learning for classification and its convergence
Merad, Ibrahim, Yu, Yiyang, Bacry, Emmanuel, Gaïffas, Stéphane
The aim of this work is to provide additional theoretical guarantees for contrastive learning (van den Oord et al., 2018), which corresponds to methods allowing to learn useful data representations in an unsupervised setting. Unsupervised representation learning was initially approached with a fair amount of success by training through the minimization of losses coming from "pretext" tasks, a technique known as self-supervision (Doersch and Zisserman, 2017), where labels can be automatically constructed. Notable examples of pretext tasks in computer vision include colorization (Zhang et al., 2016), transformation prediction (Gidaris et al., 2018; Dosovitskiy et al., 2014) or predicting patch relative positions (Doersch et al., 2015). Some theoretical guarantees (Lee et al., 2020) were recently proposed to support training on pretext tasks. Contrastive learning is also known to be very effective for pretraining supervised methods (Chen et al., 2020a,b; Grill et al., 2020; Caron et al., 2020), where we can observe that, quite surprisingly, the gap between unsupervised and supervised performance has been closed for tasks such as image classification: the use of a pretrained image encoder on top of simple classification layers, that are trained on a fraction of the labels available, allows to achieve an accuracy comparable to that of a fully supervised end-to-end training (Hénaff et al., 2019; Grill et al., 2020).
Dec-2-2020
- Country:
- Genre:
- Research Report (0.51)
- Technology: