r/MachineLearning - [D] Deep Learning optimization

#artificialintelligence 

I'm going to do a comparison on recent (or at least lesser-known) gradient optimization methods. The ones I've encountered up to now are the following: I would be particularly interested in approaches not using any hyperparameters at all (such as number 3 - COCOB), however I will consider all of the interesting and promising methods. Are you aware of some novelties or lesser-known approaches? Previously I've posted the question in r/MLQuestions but haven't received any feedback so I'm posting it here, hope it's not violating any rules.