A Sample Complexity Separation between Non-Convex and Convex Meta-Learning
Saunshi, Nikunj, Zhang, Yi, Khodak, Mikhail, Arora, Sanjeev
We consider the problem of meta-learning, or learning-to-learn [Thrun and Pratt, 1998], in which the goal is to use the data from numerous training tasks to reduce the sample complexity of an unseen but related test task. Although there is a long history of successful methods in meta-learning and the related areas of multi-task and lifelong learning [Evgeniou and Pontil, 2004, Ruvolo and Eaton, 2013], recent approaches have been developed with the diversity and scale of modern applications in mind. This has given rise to simple, model-agnostic methods that focus on learning a good initialization for some gradient-based method such as stochastic gradient descent (SGD), to be run on samples from a new task [Finn et al., 2017, Nichol et al., 2018]. These methods have found widespread applications in a variety of areas such as computer vision [Nichol et al., 2018], reinforcement learning [Finn et al., 2017], and federated learning [McMahan et al., 2017]. Inspired by their popularity, several recent learning-theoretic analyses of meta-learning have followed suit, eschewing customization to specific hypothesis classes such as halfspaces [Maurer and Pontil, 2013, Balcan et al., 2015] and instead favoring the convex-case study of gradient-based algorithms that could potentially be applied to deep neural networks [Denevi et al., 2019, Khodak et al., 2019].
Feb-25-2020
- Country:
- North America > United States > Pennsylvania > Allegheny County > Pittsburgh (0.04)
- Genre:
- Research Report (0.82)
- Industry:
- Education (0.34)
- Technology: