Forboundednonconvex losses and a batch sizem = 1, we additionally show that both generalization error and learning rate are independent ofd and K, and remain essentially the same asfortheSGD, evenfortwofunction evaluations.
Consider the problem of training deep neural networks on large annotated datasets, such as ImageNet [1]. This problem can be formalized as finding optimal parameters for a given neural networka,parameterized byw,w.r.t.