Joint Embedding Self-Supervised Learning in the Kernel Regime
Kiani, Bobak T., Balestriero, Randall, Chen, Yubei, Lloyd, Seth, LeCun, Yann
–arXiv.org Artificial Intelligence
In this setting, we explore two different data-augmentation (DA) policies, one aligned with the data distribution (rotation+translation+scaling) and one largely misaligned with the data distribution (aggressive Gaussian blur). Because our goal is to understand how much DA impacts the SSL kernel compared to a fully supervised benchmark, we consider two (supervised) benchmarks: one that employs the labels of the sampled training set and all the augmented samples and one that only employs the sampled training set and no augmented samples. We explore a small training set size going from N = 16 to N = 256 and for each case we produce a number of augmented samples for each datapoint so that the total number of samples does not exceed 50, 000 which is a standard threshold for kernel methods. We note that our implementation is based on the common Python scientific library Numpy/Scipy (Harris et al., 2020) and runs on CPU. We observed in Figure 3 that the SSL kernel is able to match and even outperform the fully supervised case when employing the correct data-augmentation, while with the incorrect data-augmentation, the performance is not even able to match the supervised case that did not see the augmented samples. To better understand the impact of different hyper-parameters onto the two kernels, we also study in Figure 4 the MNIST test set performances when varying the representation dimension K, the SVM's l
arXiv.org Artificial Intelligence
Sep-29-2022
- Country:
- North America > United States
- Michigan (0.04)
- Massachusetts > Middlesex County
- Cambridge (0.14)
- Europe > United Kingdom
- England > Cambridgeshire > Cambridge (0.04)
- North America > United States
- Genre:
- Research Report > New Finding (0.92)
- Technology: