Non-asymptotic approximations of neural networks by Gaussian processes

Eldan, Ronen, Mikulincer, Dan, Schramm, Tselil

arXiv.org Machine Learning 

In the past decade, artificial neural networks have experienced an unprecedented renaissance. However, the current theory has yet to catch-up with the practice and cannot explain their impressive performance. Particularly intriguing is the fact that over-parameterized models do not tend to over-fit, even when trained to zero error on the training set. Owing to this seemingly paradoxical fact, researchers have focused on understanding the infinite-width limit of neural networks. This line of research has led to many important discoveries such as the'lazy-training' regime [9,32] which is governed by the limiting'neural tangent kernel' (see [19]), as well as the'mean-field' limit approach (see [22,25,26] for some examples) to study the training dynamics and loss landscape. The first to study the limiting distribution of a neural network at (a random) initialization was Neal [28], who proved a Central Limit Theorem (CLT) for two-layered wide neural networks. According to Neal's result, when initialized with random weights, as the width of the network goes to infinity, its law converges, in distribution, to a Gaussian process. Subsequent works have generalized this result to deeper networks and other architectures ([13, 17, 24, 30, 33, 35, 36]). This correspondence between Gaussian processes and neural networks has proved to be highly influential and has inspired many new models (see [35] for a thorough review of these models).

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found