Review for NeurIPS paper: Reconciling Modern Deep Learning with Traditional Optimization Analyses: The Intrinsic Learning Rate
–Neural Information Processing Systems
Weaknesses: Among several, your paper makes two concrete predictions: 1. When dropping learning rate by 10, the intrinsic learning rate drops by 10 immediately (this is obvious), but it eventually converges to sqrt(10) 2. Reaching equilibrium takes O(1/\lambda_e) steps. I'd like to see experiments measuring and verifying them, or if your results are already in the paper, have them be more prominent, and linked to where these predictions are discussed. For example, I'd like to see a plot that plots 1/\lambda_e vs "step to convergence", which should be linear if your prediction is correct. Other questions I have 1.
Neural Information Processing Systems
Jan-27-2025, 08:40:46 GMT
- Technology: