Contrasting the landscape of contrastive and non-contrastive learning
Pokle, Ashwini, Tian, Jinjin, Li, Yuchen, Risteski, Andrej
Recent improvements in representation learning without supervision were driven by self-supervised learning approaches, in particular contrastive learning (CL), which constructs positive and negative samples out of unlabeled dataset via data augmentation (Chen et al., 2020; He et al., 2020; Caron et al., 2020; Ye et al., 2019; Oord et al., 2018; Wu et al., 2018). Subsequent works based on data augmentation also showed promising results for methods based on non-contrastive learning (non-CL), which do not require explicit negative samples (Grill et al., 2020; Richemond et al., 2020; Chen & He, 2021; Zbontar et al., 2021; Tian et al., 2021). However, understanding of how these approaches work, especially of how the learned representations compare--qualitatively and quantitatively--is lagging behind. In this paper, via a combination of empirical and theoretical results, we provide evidence that non-contrastive methods based on data augmentation can lead to substantially worse representations. Most notably, avoiding the collapsed representations has been the key ingredient in prior successes in non-contrastive learning. The collapses was first referred as the complete collapse, that is, all representation vectors shrink into a single point; later a new type of collapses, dimension collapse (Hua et al., 2021; Jing et al., 2022) caught attention as well, that is the embedding vectors only span a lower-dimensional subspace.
Mar-29-2022
- Country:
- North America > United States
- Pennsylvania > Allegheny County > Pittsburgh (0.04)
- Europe > France
- Île-de-France > Paris > Paris (0.04)
- North America > United States
- Genre:
- Research Report (1.00)
- Technology: