A geometric interpretation of stochastic gradient descent using diffusion metrics
Fioresi, R., Chaudhari, P., Soatto, S.
Stochastic gradient descent (SGD) is a key ingredient in the training of deep neural networks and yet its geometrical significance appears elusive. We study a deterministic model in which the trajectories of our dynamical systems are described via geodesics of a family of metrics arising from the diffusion matrix. These metrics encode information about the highly non-isotropic gradient noise in SGD. We establish a parallel with General Relativity models, where the role of the electromagnetic field is played by the gradient of the loss function. We compute an example of a two layer network.
Oct-27-2019
- Country:
- North America > United States
- Pennsylvania (0.04)
- New York (0.04)
- California (0.04)
- Europe > Italy
- Emilia-Romagna > Metropolitan City of Bologna > Bologna (0.05)
- North America > United States
- Genre:
- Research Report (0.50)
- Technology: