The Surprising Simplicity of the Early-Time Learning Dynamics of Neural Networks
Hu, Wei, Xiao, Lechao, Adlam, Ben, Pennington, Jeffrey
Modern deep learning models are enormously complex function approximators, with many state-of-the-art architectures employing millions or even billions of trainable parameters [Radford et al., 2019, Adiwardana et al., 2020]. While the raw parameter count provides only a crude approximation of a model's capacity, more sophisticated metrics such as those based on PAC-Bayes [McAllester, 1999, Dziugaite and Roy, 2017, Neyshabur et al., 2017b], VC dimension [Vapnik and Chervonenkis, 1971], and parameter norms [Bartlett et al., 2017, Neyshabur et al., 2017a] also suggest that modern architectures have very large capacity. Moreover, from the empirical perspective, practical models are flexible enough to perfectly fit the training data, even if the labels are pure noise [Zhang et al., 2017]. Surprisingly, these same high-capacity models generalize well when trained on real data, even without any explicit control of capacity. These observations are in conflict with classical generalization theory, which contends that models of intermediate complexity should generalize best, striking a balance between the bias and the variance of their predictive functions. To reconcile theory with observation, it has been suggested that deep neural networks may enjoy some form of implicit regularization induced by gradient-based training algorithms that biases the trained models towards simpler functions.
Jun-25-2020
- Country:
- Europe > United Kingdom > England > Cambridgeshire > Cambridge (0.04)
- Genre:
- Research Report (0.82)
- Technology: