Analysis and Visualization of Classical and Data-Based Neural-Network Initializations

#artificialintelligence 

Fundamental breakthroughs have been critical to the growth of deep learning. Almost all networks can benefit from the ensembling effect of residual connections [Veit et al., 2016], like-wise many networks can be improved by the effects (such as regularization [De and Smith, 2020, Hoffer et al., 2018]) afforded by Batch Normalization. This brings us to our present focus, neural network initialization; all networks can benefit from better initialization. One of the earliest works on initialization is by Glorot and Bengio [Glorot and Bengio, 2010]. Before this work, it was difficult to train deep networks at all, as the theory behind training such networks had not been adequately explored. Thus, deep networks were prone to the exploding or vanishing gradient problem; activations would either explode or vanish as they traveled down the network, and thus gradients would too, making learning impossible [Glorot and Bengio, 2010]. As explained, this was initially solved in part by Glorot and Bengio, who propose a theoretically sound initialization method for a number of symmetric activation functions. This result was later extended by way of Kaiming initialization [He et al., 2015b] to the ReLU initialization, allowing the training of a class of deep and efficient ReLU networks. We consider these more basic and theoretically-based initializations classical neural network initialization methods.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found