AITopics | stability and generalization analysis

Collaborating Authors

stability and generalization analysis

Information about AI from the News, Publications, and Conferences

Automatic Classification – Tagging and Summarization – Customizable Filtering and Analysis

If you are looking for an answer to the question What is Artificial Intelligence? and you only have a minute, then here's the definition the Association for the Advancement of Artificial Intelligence offers on its home page: "the scientific understanding of the mechanisms underlying thought and intelligent behavior and their embodiment in machines."

However, if you are fortunate enough to have more than a minute, then please get ready to embark upon an exciting journey exploring AI (but beware, it could last a lifetime) …

Stability and Generalization Analysis of Gradient Methods for Shallow Neural Networks

Neural Information Processing SystemsDec-25-2025, 19:11:23 GMT

While significant theoretical progress has been achieved, unveiling the generalization mystery of overparameterized neural networks still remains largely elusive. In this paper, we study the generalization behavior of shallow neural networks (SNNs) by leveraging the concept of algorithmic stability. We consider gradient descent (GD) and stochastic gradient descent (SGD) to train SNNs, for both of which we develop consistent excess risk bounds by balancing the optimization and generalization via early-stopping. As compared to existing analysis on GD, our new analysis requires a relaxed overparameterization assumption and also applies to SGD. The key for the improvement is a better estimation of the smallest eigenvalues of the Hessian matrices of the empirical risks and the loss function along the trajectories of GD and SGD by providing a refined estimation of their iterates.

gradient method, name change, stability and generalization analysis, (5 more...)

Neural Information Processing Systems

Technology:

Information Technology > Artificial Intelligence > Machine Learning > Statistical Learning > Gradient Descent (0.86)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks (0.56)

Add feedback

Appendix for " Stability and Generalization Analysis of Gradient Methods for Shallow Neural Networks " A Lemmas

Neural Information Processing SystemsAug-19-2025, 21:46:32 GMT

In this section, we collect several lemmas useful for our analysis. The proof is completed.Lemma A.2. Let W, W According to Taylor's theorem, there exists α [0, 1] such that ℓ(W; z) ℓ(W The proof is completed.The following lemma shows the self-bounding property of smooth and nonnegative functions. Let Assumptions 1, 2 hold. Let Assumptions 1, 2 hold. The remaining arguments in proving Lemma A.6 is the same as proving Lemma 5 in [ Let Assumptions 1, 2 hold.

artificial intelligence, lemma, machine learning, (15 more...)

Neural Information Processing Systems

Technology: Information Technology > Artificial Intelligence > Machine Learning > Neural Networks (0.40)

Add feedback

Stability and Generalization Analysis of Gradient Methods for Shallow Neural Networks

Neural Information Processing SystemsJan-19-2025, 08:10:03 GMT

gradient method, shallow neural network, stability and generalization analysis, (3 more...)

Neural Information Processing Systems

Technology:

Information Technology > Artificial Intelligence > Machine Learning > Statistical Learning > Gradient Descent (0.93)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks (0.93)

Add feedback