Statistical Learning
Few-Shot Data-Driven Algorithms for Low Rank Approximation
Recently, data-driven and learning-based algorithms for low rank matrix approximation were shown to outperform classical data-oblivious algorithms by wide margins in terms of accuracy. Those algorithms are based on the optimization of sparse sketching matrices, which lead to large savings in time and memory during testing. However, they require long training times on a large amount of existing data, and rely on access to specialized hardware and software. In this work, we develop new data-driven low rank approximation algorithms with better computational efficiency in the training phase, alleviating these drawbacks. Furthermore, our methods are interpretable: while previous algorithms choose the sketching matrix either at random or by black-box learning, we show that it can be set (or initialized) to clearly interpretable values extracted from the dataset. Our experiments show that our algorithms, either by themselves or in combination with previous methods, achieve significant empirical advantages over previous work, improving training times by up to an order of magnitude toward achieving the same target accuracy.
Improving Self-Supervised Learning by Characterizing Idealized Representations
Our goal is to provide a simple conceptual framework to think about those questions. To derive such a framework, we ask ourselves: what are the ideal requirements that ISSL representations should aim to satisfy? We prove necessary and sufficient requirements to ensure that probes from a specified family, e.g.