Technology
d9731321ef4e063ebbee79298fa36f56-AuthorFeedback.pdf
Our analysis provides full distribution information on the joint outputs. Furthermore, the9 distribution ofthe cosine similarity explains whymoderately deepand wide ReLU networks can betrained despite10 negative results by mean field (MF) analysis based on correlations. There,14 the normal distribution originates from the MF limit. In contrast, here we understand that the output distribution is15 completely determined bytheempirical covariance matrix ofinputs. This is rather obvious however. Instead, we refer to the rich literature on linear neural networks at23 initialization.
d921c3c762b1522c475ac8fc0811bb0f-AuthorFeedback.pdf
We wish to thank all of the reviewers for their time and thorough reading of our paper! We appreciate the reviewer's suggestions regarding clarity. We have added the suggested summary sentence "the key We started with binary sentiment classification, but are actively working on more tasks. RNN hidden states onto the top two PCs for two different input sequences that differ only by two tokens (replacing ' The trajectories start out the same as the initial tokens are identical. We have added a footnote noting this in the main text.