Goto

Collaborating Authors

 Country



TheGeneralization-StabilityTradeoffInNeural NetworkPruning

Neural Information Processing Systems

This motivation is particularly relevant given the perhaps surprising observation that a wide variety of pruning approaches increase test accuracy despite sometimes massivereductions inparameter counts.






In the following experiments, unless otherwise explicitly stated,we use the DMGC model asthe graphmatchingmethodinthissection. (a) Metastepsizeβ (b) # SamplesN (c) Initialbandwidthb0 (d) Parameters

Neural Information Processing Systems

We also include equal number of randomly selected disconnected links that servers as negative samples. The performance curves initially raise and then drop quickly whenβ continuously increases. This demonstrates that there must exist the optimalβ that makes the meta learning be maximally 20 optimized. Sensitivity of number of samples N. Figure 7 (b) exhibits the sensitivity ofN in our MLPGD model withN between 1 and 15. All models were trained for 500 iterations, with a batch size of 512, and a learningrateof0.001.




Knowledge Distillation in Wide Neural Networks: Risk Bound, Data Efficiency and Imperfect Teacher

Neural Information Processing Systems

On the other hand, recent finding on neural tangent kernel enables us to approximate a wide neural network with a linear model of the network's random features. In this paper, we theoretically analyze the knowledge distillation of a wide neural network. First we provide a transfer risk bound for the linearized model of the network. Then we propose a metric of the task's training difficulty, called data inefficiency.