blog.md
Despite wide adoption in the industry, our understanding of deep learning is still lagging. Theorists have long assumed networks with hundreds of thousands of neurons and orders of magnitude more individually weighted connections between them should suffer from a fundamental problem: over-parameterization [19] The spin glass model seems a great starting point for a very fertile research direction. Should we go down this path? First, the ultimate metric in machine learning is the generalisation accuracy rather than minimising the loss function, which only captures how well the model fits the training data [...] Second, it is known empirically that deeper minima have lower generalisation accuracy than shallower minima Before continuing our work on loss function topology, let's have a look at what is meant with flatness of minima. Consider the cartoon energy landscape (of the empirical loss function) in figure X discussed in [5].
Oct-13-2018, 23:23:58 GMT
- Technology: