d77f00766fd3be3f2189c843a6af3fb2-AuthorFeedback.pdf

Neural Information Processing Systems 

This part was poorly presented; we have updated the text[Update B]. To15 measure training instead of training confounded with issues of memorization vs. generalization, we use thetrain16 gradients instead ofval. Observations like the last layer hurting are also more surprising on train vs. val. We use17 full-batchgradients for analysis instead of single mini-batch gradient to measure learning in as noise-free a way as18 possible. First order was only mentioned as an example to illustrate the concept.22

Similar Docs  Excel Report  more

TitleSimilaritySource
None found