Why to "grow" and "harvest" deep learning models?
Kulikovskikh, Ilona, Legović, Tarzan
Current expectations from training deep learning models with gradient-based methods include: 1) transparency; 2) high convergence rates; 3) high inductive biases. While the state-of-art methods with adaptive learning rate schedules are fast, they still fail to meet the other two requirements. We suggest reconsidering neural network models in terms of single-species population dynamics where adaptation comes naturally from open-ended processes of "growth" and "harvesting". We show that the stochastic gradient descent (SGD) with two balanced pre-defined values of per capita growth and harvesting rates outperform the most common adaptive gradient methods in all of the three requirements.
Aug-8-2020
- Country:
- Asia > Russia (0.04)
- North America > United States
- New York > New York County
- New York City (0.04)
- Massachusetts > Middlesex County
- Cambridge (0.04)
- New York > New York County
- Europe
- United Kingdom > England
- Cambridgeshire > Cambridge (0.14)
- Russia > Volga Federal District
- Samara Oblast > Samara (0.04)
- Croatia > Zagreb County
- Zagreb (0.04)
- United Kingdom > England
- Genre:
- Research Report (0.64)
- Technology: