Goto

Collaborating Authors

 Deep Learning



CD_GraB_camera_ready

Neural Information Processing Systems

Whereas RR arbitrarily permutes training examples, GraB leverages stale gradients from prior epochs to order examples -- achieving a provably faster convergence rate than RR.






Provable Guarantees for Neural Networks via Gradient Feature Learning

Neural Information Processing Systems

Neural networks have achieved remarkable empirical performance, while the current theoretical analysis is not adequate for understanding their success, e.g., the Neural Tangent Kernel approach fails to capture their key feature learning ability, while recent analyses on feature learning are typically problem-specific.