Size-free generalization bounds for convolutional neural networks
Long, Philip M., Sedghi, Hanie
–arXiv.org Artificial Intelligence
Recently, substantial progress has been made regarding theoretical analysis of the generalization of deep learning models [see Zhang et al., 2016, Dziugaite and Roy, 2017, Bartlett et al., 2017, Neyshabur et al., 2017, 2018, Arora et al., 2018, Neyshabur et al., 2019]. One interesting point that has been explored, with roots in [Bartlett, 1998], is that even if there are many parameters, the set of models computable using weights with small magnitude is limited enough to provide leverage for induction [Bartlett et al., 2017, Neyshabur et al., 2018]. Intuitively, if the weights start small, since the most popular training algorithms make small, incremental updates that get smaller as the training accuracy improves, there is a tendency for these algorithms to produce small weights.
arXiv.org Artificial Intelligence
Jun-12-2019