The Case for Full-Matrix Adaptive Regularization
Agarwal, Naman, Bullins, Brian, Chen, Xinyi, Hazan, Elad, Singh, Karan, Zhang, Cyril, Zhang, Yi
Stochastic gradient descent is the workhorse behind the recent deep learning revolution. This simple and ageold algorithm has been supplemented with a variety of enhancements to improve its practical performance, and sometimes its theoretical guarantees. Amongst the acceleration methods there are three main categories: momentum, adaptive regularization, and variance reduction. Momentum (in its various incarnations, like heavy-ball or Nesterov acceleration) is the oldest enhancement. It has a well-developed theory, and is known to improve practical convergence in a variety of tasks, small and large. It is also easy to implement.
Jun-7-2018