Intro to optimization in deep learning: Busting the myth about batch normalization

Open in new window