Estimation of a regression function on a manifold by fully connected deep neural networks

Kohler, Michael, Langer, Sophie, Reif, Ulrich

arXiv.org Machine Learning 

Deep neural networks (DNNs) are built of multiple layers and learn sequentially multiple levels of representation and abstraction by performing a nonlinear transformation on the data. The approach has proven itself to work incredibly well in practice, like for speech (Graves et al. (2013)) and image recognition (Krizhevsky et al. (2017)), or game intelligence (Silver et al. (2016)). But, unfortunately, the procedure is not well understood. Recently, several researchers tried to explain the performance of DNNs from a theoretical point of view. Results concerning the approximation power of DNNs were shown in Montufar (2014), Eldan and Shamir (2016), Yarotsky (2017), Yarotsky and Zhevnerchuck (2020), Langer (2021b) and Lu et al. (2020). Beside this, quite a few articles try to answer the question about why neural networks perform well on unknown new data sets (cf., e.g., Bauer and Kohler (2019), Schmidt-Hieber (2020), Kohler and Langer (2020), Kohler, Krzyżak, and Langer (2019), Langer (2021a), Imaizumi and Fukumizu