Calibrated Chaos: Variance Between Runs of Neural Network Training is Harmless and Inevitable

Jordan, Keller

arXiv.org Artificial Intelligence 

Modern neural networks (Krizhevsky et al., 2012; He et al., 2016; Vaswani et al., 2017) are trained using stochastic gradient-based algorithms (Rumelhart et al., 1986; Kingma and Ba, 2014), involving randomized weight initialization, data/batch ordering, and data augmentations. Because of this stochasticity, each independent run of training produces a different network with better or worse performance than average. The difference between such independent runs is often substantial. Picard (2021) finds that for a standard CIFAR-10 (Krizhevsky et al., 2009) training configuration, there exist random seeds which differ by 1.3% in terms of test-set accuracy. In comparison, the gap between the top two methods competing for state-of-the-art on CIFAR-10 has been less than 1% throughout the majority of the benchmark's lifetime

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found