Stepping on the Edge: Curvature A ware Learning Rate Tuners

Neural Information Processing Systems 

(Liu and Nocedal, 1989). Similar efforts have been made for Polyak stepsizes (Berrada et al., 2020; Loizou et al., 2021), in addition to new methods which combine distance to optimality with online learning convergence bounds (Cutkosky et al., 2023; Classically-inspired methods, however, have generally struggled to gain traction in deep learning.

Similar Docs  Excel Report  more

TitleSimilaritySource
None found