Overcoming Multi-Model Forgetting
Benyahia, Yassine, Yu, Kaicheng, Bennani-Smires, Kamil, Jaggi, Martin, Davison, Anthony, Salzmann, Mathieu, Musat, Claudiu
When We identify a phenomenon, which we refer to dealing with many large models, a common strategy to keep as multi-model forgetting, that occurs when sequentially training tractable is to share a subset of the weights across training multiple deep networks with the multiple models and to train them sequentially (Pham partially-shared parameters; the performance of et al., 2018; Xie & Yuille, 2017; Liu et al., 2018a). This previously-trained models degrades as one optimizes strategy has a major drawback. Figure 1 shows that for two a subsequent one, due to the overwriting models, A and B, the larger the number of shared weights, of shared parameters. To overcome this, we introduce the more the accuracy of A drops when training B; B overwrites a statistically-justified weight plasticity loss some of the weights of A and this damages the performance that regularizes the learning of a model's shared of A. We call this multi-model forgetting. The parameters according to their importance for the benefits of weight-sharing have been emphasized in tasks previous models, and demonstrate its effectiveness like neural architecture search, where the associated speed when training two models sequentially and gains have been key in making the process practical (Pham for neural architecture search. Adding weight et al., 2018; Liu et al., 2018b), but its downsides remain plasticity in neural architecture search preserves virtually unexplored.
Mar-2-2019