Harmless Overparametrization in Two-layer Neural Networks

Wang, Huiyuan, Lin, Wei

arXiv.org Machine Learning 

In such a wide range of applications, neural networks are preferred to be overparametrized in the sense that the number of active (nonzero) network parameters is much larger than the sample size. One example is the Alex-Net (Krizhevsky, Sutskever and Hinton, 2012). There is convincing evidence showing that overparametrization can help optimization (Arora, Cohen and Hazan, 2018; Safran, Yehudai and Shamir, 2020), however, overparametrized deep neural networks can easily fit random labels even in the presence of explicit regularization (Zhang et al., 2016). This indicates that overparametrized models have the capability to overfit, but do not necessarily lead to bad testing performance. Thus two intriguing questions come out: why can overparametrized neural networks exhibit good testing performance and how good can the testing performance be? Overparametrization, where the number of active parameters is larger than the sample size, is not new in statistics. For example, high dimensional linear models can be viewed to be overparametrized. In high dimensional linear regression, prediction is closely related to the estimation of the true parameter and overparametrization can be harmful to estimation even in the presence of optimal regularization. For example, the prediction risk of the optimal regularized ridge estimator can have a strictly positive limit in the overparametrized settings (Dobriban et al., 2018; Hastie et al., 2019).

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found