generalized neural tangent kernel analysis
A Generalized Neural Tangent Kernel Analysis for Two-layer Neural Networks
A recent breakthrough in deep learning theory shows that the training of over-parameterized deep neural networks can be characterized by a kernel function called \textit{neural tangent kernel} (NTK). However, it is known that this type of results does not perfectly match the practice, as NTK-based analysis requires the network weights to stay very close to their initialization throughout training, and cannot handle regularizers or gradient noises. In this paper, we provide a generalized neural tangent kernel analysis and show that noisy gradient descent with weight decay can still exhibit a ``kernel-like'' behavior. This implies that the training loss converges linearly up to a certain accuracy. We also establish a novel generalization error bound for two-layer neural networks trained by noisy gradient descent with weight decay.
Review for NeurIPS paper: A Generalized Neural Tangent Kernel Analysis for Two-layer Neural Networks
Weaknesses: The authors claim without justification that the standard tools used in NTK regimes CANNOT handle noisy gradient and regularizers (Line 28 and 168). It is not clear that whether the prior works just did not analyze these two particular scenarios due to limited space or similar, or these two scenarios cannot be handled in principle. If it is the first case, the results in this paper would be not quite significant, and can be considered as a natural extension of previous works. So, I suggest the author to spend a section or so to provide detailed analysis on how and why noisy gradient and regularizers can not be handled by standard NTK analysis. Note that the model f contains a scaling factor alpha, see Eq.(3.1) and (3.2).
Review for NeurIPS paper: A Generalized Neural Tangent Kernel Analysis for Two-layer Neural Networks
This paper extends neural tangent kernel results to a two-layer, infinite width neural network with a three-times differentiable activation function, weight decay regularization, and noisy gradient descent training, showing a linear convergence rate. The paper received mixed reviews (marginally above, marginally below, accept, reject). On the positive side, R3 think the results are a new nontrivial extension of the NTK results, and R1 think the paper is novel, well written, etc. R1 had some technical issues, but was satisfied by the rebuttal. On the other hand, R2 raised some technical issues regarding the effect of the scaling in the kernel, which I think are well addressed by the rebuttal. R4's main critique is that he/she is not convinced about the significance of using L2 regularization, since algorithms have implicit regularization.
A Generalized Neural Tangent Kernel Analysis for Two-layer Neural Networks
A recent breakthrough in deep learning theory shows that the training of over-parameterized deep neural networks can be characterized by a kernel function called \textit{neural tangent kernel} (NTK). However, it is known that this type of results does not perfectly match the practice, as NTK-based analysis requires the network weights to stay very close to their initialization throughout training, and cannot handle regularizers or gradient noises. In this paper, we provide a generalized neural tangent kernel analysis and show that noisy gradient descent with weight decay can still exhibit a kernel-like'' behavior. This implies that the training loss converges linearly up to a certain accuracy. We also establish a novel generalization error bound for two-layer neural networks trained by noisy gradient descent with weight decay.