On Dissipativity of Cross-Entropy Loss in Training ResNets
Püttschneider, Jens, Faulwasser, Timm
–arXiv.org Artificial Intelligence
For example, the backpropagation in neural network training appears in optimal control as adjoint sensitivity equation (Esteve-Yagüe and Geshkovski, 2023; Faulwasser et al., 2021; Esteve et al., 2021). Moreover, the training of Neural Networks (NNs) with constant width in each layer can be formulated as an Optimal Control Problem (OCP) (Li et al., 2018; Esteve et al., 2021). In this context, the layer-to-layer propagation of the data is considered a dynamical system on a finite horizon corresponding to the depth of the network. The solution to the OCP determines the weights and biases of the neural network, i.e., weights and biases are the control inputs to drive the data to a desired point in the terminal layer determined by the label and loss function. In particular, the system and control perspective is helpful for Residual Neural Networks (ResNets) (He et al., 2016), which can be regarded as Euler forward discretizations of neural ODEs (Chen et al., 2018). Chang et al., 2018 analyze the reversibility and stability of the ResNet dynamics based on their continuous time counterparts. Esteve et al., 2021 and Faulwasser et al., 2021 have suggested to include a regularization term based on the states of the hidden layers in the training OCP. Faulwasser et al., 2021 analyze the dissipativity and the related turnpike property from an optimal control point of view when utilizing a quadratic (l
arXiv.org Artificial Intelligence
May-29-2024
- Country:
- North America > United States
- New York (0.04)
- Europe
- Austria > Vienna (0.14)
- United Kingdom > England
- Cambridgeshire > Cambridge (0.04)
- Germany
- Hamburg (0.04)
- North Rhine-Westphalia > Arnsberg Region
- Dortmund (0.04)
- North America > United States
- Genre:
- Research Report (0.82)
- Industry:
- Health & Medicine (0.95)
- Energy (0.68)
- Technology: