GD doesn't make the cut: Three ways that non-differentiability affects neural network training

Open in new window