Extracurricular Learning: Knowledge Transfer Beyond Empirical Distribution

Pouransari, Hadi, Tuzel, Oncel

arXiv.org Machine Learning 

For example, both the PyramidNet-110 model [23] and the larger PyramidNet-Knowledge distillation has been used to transfer 200 model achieve perfect accuracy on the CIFAR100 [32] knowledge learned by a sophisticated model (teacher) to training set, while the latter has 3% higher generalization a simpler model (student). This technique is widely used to accuracy. This motivated transferring the "knowledge" compress model complexity. However, in most applications encoded in the more accurate larger model to the smaller the compressed student model suffers from an accuracy gap one. Knowledge Distillation [8, 27] (KD) established with its teacher. We propose extracurricular learning, a an important mechanism through which one model novel knowledge distillation method, that bridges this gap (typically of higher capacity, called teacher) can train by (1) modeling student and teacher output distributions; another model (typically a smaller model that satisfies (2) sampling examples from an approximation to the the computational budget, called student). KD has been underlying data distribution; and (3) matching student and implemented in many machine learning tasks, for example teacher output distributions over this extended set including image classification [27], object detection [12, 65], video uncertain samples. We conduct rigorous evaluations on labeling [74], natural language processing [60, 41, 57, 36, regression and classification tasks and show that compared 61], and speech recognition [11, 59, 37]. to the standard knowledge distillation, extracurricular The idea of KD is to encourage the student to imitate learning reduces the gap by 46% to 68%. This leads to teacher's behavior over a set of data points, called transferset.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found