Joint Embedding Self-Supervised Learning in the Kernel Regime

Kiani, Bobak T., Balestriero, Randall, Chen, Yubei, Lloyd, Seth, LeCun, Yann

arXiv.org Artificial Intelligence 

In this setting, we explore two different data-augmentation (DA) policies, one aligned with the data distribution (rotation+translation+scaling) and one largely misaligned with the data distribution (aggressive Gaussian blur). Because our goal is to understand how much DA impacts the SSL kernel compared to a fully supervised benchmark, we consider two (supervised) benchmarks: one that employs the labels of the sampled training set and all the augmented samples and one that only employs the sampled training set and no augmented samples. We explore a small training set size going from N = 16 to N = 256 and for each case we produce a number of augmented samples for each datapoint so that the total number of samples does not exceed 50, 000 which is a standard threshold for kernel methods. We note that our implementation is based on the common Python scientific library Numpy/Scipy (Harris et al., 2020) and runs on CPU. We observed in Figure 3 that the SSL kernel is able to match and even outperform the fully supervised case when employing the correct data-augmentation, while with the incorrect data-augmentation, the performance is not even able to match the supervised case that did not see the augmented samples. To better understand the impact of different hyper-parameters onto the two kernels, we also study in Figure 4 the MNIST test set performances when varying the representation dimension K, the SVM's l

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found