Goto

Collaborating Authors

 neural network estimate


Statistically guided deep learning

arXiv.org Machine Learning

We present a theoretically well-founded deep learning algorithm for nonparametric regression. It uses over-parametrized deep neural networks with logistic activation function, which are fitted to the given data via gradient descent. We propose a special topology of these networks, a special random initialization of the weights, and a data-dependent choice of the learning rate and the number of gradient descent steps. We prove a theoretical bound on the expected $L_2$ error of this estimate, and illustrate its finite sample size performance by applying it to simulated data. Our results show that a theoretical analysis of deep learning which takes into account simultaneously optimization, generalization and approximation can result in a new deep learning estimate which has an improved finite sample performance.


Analysis of the expected $L_2$ error of an over-parametrized deep neural network estimate learned by gradient descent without regularization

arXiv.org Machine Learning

Recent results show that estimates defined by over-parametrized deep neural networks learned by applying gradient descent to a regularized empirical $L_2$ risk are universally consistent and achieve good rates of convergence. In this paper, we show that the regularization term is not necessary to obtain similar results. In the case of a suitably chosen initialization of the network, a suitable number of gradient descent steps, and a suitable step size we show that an estimate without a regularization term is universally consistent for bounded predictor variables. Additionally, we show that if the regression function is H\"older smooth with H\"older exponent $1/2 \leq p \leq 1$, the $L_2$ error converges to zero with a convergence rate of approximately $n^{-1/(1+d)}$. Furthermore, in case of an interaction model, where the regression function consists of a sum of H\"older smooth functions with $d^*$ components, a rate of convergence is derived which does not depend on the input dimension $d$.


A Free Lunch with Influence Functions? Improving Neural Network Estimates with Concepts from Semiparametric Statistics

arXiv.org Machine Learning

Parameter estimation in empirical fields is usually undertaken using parametric models, and such models readily facilitate statistical inference. Unfortunately, they are unlikely to be sufficiently flexible to be able to adequately model real-world phenomena, and may yield biased estimates. Conversely, non-parametric approaches are flexible but do not readily facilitate statistical inference and may still exhibit residual bias. We explore the potential for Influence Functions (IFs) to (a) improve initial estimators without needing more data (b) increase model robustness and (c) facilitate statistical inference. We begin with a broad introduction to IFs, and propose a neural network method 'MultiNet', which seeks the diversity of an ensemble using a single architecture. We also introduce variants on the IF update step which we call 'MultiStep', and provide a comprehensive evaluation of different approaches. The improvements are found to be dataset dependent, indicating an interaction between the methods used and nature of the data generating process. Our experiments highlight the need for practitioners to check the consistency of their findings, potentially by undertaking multiple analyses with different combinations of estimators. We also show that it is possible to improve existing neural networks for `free', without needing more data, and without needing to retrain them.


On the rate of convergence of a deep recurrent neural network estimate in a regression problem with dependent data

arXiv.org Machine Learning

Motivated by the huge success of deep neural networks in applications (see, e.g., Schmidhuber (2015), Rawat and Wang (2017), Hewamalage, Bergmeir and Bandara (2020) and the literature cited therein) there is nowadays a strong interest in showing theoretical properties of such estimates. In the last years many new results concerning deep feedforward neural network estimates have been derived (cf., e.g., Eldan and Shamir (2016), Lu et al. (2020), Yarotsky (2018) and Yarotsky and Zhevnerchuk (2019) concerning approximation properties or Kohler and Krzyżak (2017), Bauer and Kohler (2019) and Schmidt-Hieber (2020) concerning statistical properties of these estimates).