exploring generalization
Reviews: Exploring Generalization in Deep Learning
Update after rebuttal: quote the rebuttal "Even this very simple variability in architecture, proves challenging to study using the complexity measures suggested. We certainly intend to study also other architectural differences." I hope the authors include the discussions (and hopefully some experiment results) about applying the proposed techniques to the comparison of models with different architectures in the final version if it get accepted. The current results are already useful steps towards understanding deep neural networks by their own, but having this results is really a great add on to this paper. Given different global minimizers of the empirical risk, can we tell which one generalize better based on the complexity measure.
Exploring Generalization in Deep Learning
Neyshabur, Behnam, Bhojanapalli, Srinadh, Mcallester, David, Srebro, Nati
With a goal of understanding what drives generalization in deep networks, we consider several recently suggested explanations, including norm-based control, sharpness and robustness. We study how these measures can ensure generalization, highlighting the importance of scale normalization, and making a connection between sharpness and PAC-Bayes theory. We then investigate how well the measures explain different observed phenomena. Papers published at the Neural Information Processing Systems Conference.