Over-Parameterized Deep Neural Networks Have No Strict Local Minima For Any Continuous Activations
Li, Dawei, Ding, Tian, Sun, Ruoyu
Recently, the application of deep neural networks [1] has led to a phenomenal success in various artificial intelligence areas, e.g., computer vision, natural language processing, and audio recognition. However, the theoretical understanding of neural networks is still limited. One of the main difficulties of analyzing neural networks is the non-convexity of the objective function, which may cause many local minima. In practice, it is observed that when the number of parameters is sufficiently large, common optimization algorithmssuch as stochastic gradient descent (SGD) can achieve small training error [2-6]. These observations are often explained by the intuition that more parameters can smooth the landscape [4,7]. Among various definitions of over-parameterization, a popular one is that the last hidden layer has more neurons than the number of training samples. Even under this assumption, it is yet unclear to what extent we can prove a rigorous result. For instance, can we prove that for any neuron activation function, every local minimum is a global minimum? If not, what exactly can we prove, and what can we not prove?
Dec-28-2018
- Country:
- North America > United States
- Michigan (0.04)
- Illinois > Champaign County
- Urbana (0.04)
- California > Alameda County
- Berkeley (0.04)
- Asia
- Middle East > Jordan (0.04)
- China > Hong Kong (0.04)
- Myanmar > Tanintharyi Region
- Dawei (0.04)
- North America > United States
- Genre:
- Research Report (0.64)
- Technology: