Review for NeurIPS paper: Rethinking Importance Weighting for Deep Learning under Distribution Shift
–Neural Information Processing Systems
Weaknesses: However, it is unclear how the experiments were done. Which version of SGD was used, and why not Adam or some better optimizer than Adam in deep learning? The authors claimed it is compatible with any model and any optimizer, but didn't show it is not limited to SGD. Moreover, why the convergence analysis is interesting? Can you show the assumptions hold even if the optimizer is limited to SGD?
Neural Information Processing Systems
Jan-26-2025, 11:13:50 GMT
- Technology: