New computational algorithms make it possible to build neural networks with many input nodes and many layers, and distinguish "deep learning" of these networks from previous work on artificial neural nets.
In order to address this limitation, we propose a novel class of functions that can characterize the loss landscape of modern deep models without requiring extensive over-parametrization and can also include saddle points.
In the literature of studying training dynamics of transformers, several simplifications are commonly adopted such as weight reparameter-ization, attention linearization, special initialization, and lazy regime.