Sharp Minima Can Generalize: A Loss Landscape Perspective On Data
Fan, Raymond, Sandlund, Bryce, Ko, Lin Myat
–arXiv.org Artificial Intelligence
Generalization in neural networks refers to the phenomenon where trained models tend to perform well on unseen test sets. This is surprising since neural networks have enough parameters to fit any dataset, even to random labels or random noise [30]. Existing guarantees on generalization require limiting model capacity, forcing low complexity solutions to fit the data [27, 3]; these bounds generally do not apply to overparameterized neural networks. This has inspired the volume hypothesis, which proposes generalization arises from the loss landscape. The hypothesis states the volume of parameter space occupied by generalization minima is significantly larger than the volumes of other minima, and thus any procedure that minimizes the training loss is likely to find good minima.
arXiv.org Artificial Intelligence
Nov-10-2025