Projected Subnetworks Scale Adaptation
Datta, Siddhartha, Shadbolt, Nigel
–arXiv.org Artificial Intelligence
Large models support great zero-shot and few-shot capabilities. However, updating these models on new tasks can break performance on previous seen tasks and their zero/few-shot unseen tasks. Our work explores how to update zero/fewshot learners such that they can maintain performance on seen/unseen tasks of previous tasks as well as new tasks. By manipulating the parameter updates of a gradient-based meta learner as the projected task-specific subnetworks, we show improvements for large models to retain seen and zero/few shot task performance in online settings. The adaptation of deep neural networks have practical importance. It enables models to adapt to varying test-time distributions, attributed to shifts in time, person, environment, etc. The more difficult adaptation cases arise when there may be no clear task boundaries, when the task was not seen during training, and only few/zero samples are available to update a model. To tackle adaptation broadly, given a base learner optimizing its inner objective with respect to its assigned task, a meta learner computes the update to the base learner such that it optimizes its outer objective across a distribution of tasks (Hospedales et al., 2021).
arXiv.org Artificial Intelligence
Jan-26-2023
- Country:
- Oceania > Australia
- New South Wales > Sydney (0.04)
- North America > United States
- California > Santa Clara County > Palo Alto (0.04)
- Europe > United Kingdom
- England > Oxfordshire > Oxford (0.14)
- Asia > Middle East
- Jordan (0.04)
- Oceania > Australia
- Genre:
- Research Report (0.41)
- Technology: