Review for NeurIPS paper: AdaBelief Optimizer: Adapting Stepsizes by the Belief in Observed Gradients
–Neural Information Processing Systems
Weaknesses: 1- The paper contains some unsubstantiated claims. For instance: * Line 145: "Although the above cases are simple, they occur frequently in deep learning, hence we expect AdaBelief to outperform Adam in general cases" and line 150 "most networks behave(s) like (the) case (in) Figure 1(b)." This statement is not substantiated. Although ReLU losses are somewhat similar to L1 loss (both are composed of two linear components), if one consider the composition resulting from several layers of a deep neural net, the resulting loss function is no longer a simple piecewise linear convex function such as the examples in Fig 3 (a), (b) and (d). That is not to stay that these examples are not interesting.
Neural Information Processing Systems
Feb-6-2025, 23:46:41 GMT
- Technology: