Benefit of deep learning with non-convex noisy gradient descent: Provable excess risk bound and superiority to kernel methods

Open in new window