dissipativity
Model Predictive Control is Almost Optimal for Restless Bandit
Gast, Nicolas, Narasimha, Dheeraj
We consider the discrete time infinite horizon average reward restless markovian bandit (RMAB) problem. We propose a \emph{model predictive control} based non-stationary policy with a rolling computational horizon $\tau$. At each time-slot, this policy solves a $\tau$ horizon linear program whose first control value is kept as a control for the RMAB. Our solution requires minimal assumptions and quantifies the loss in optimality in terms of $\tau$ and the number of arms, $N$. We show that its sub-optimality gap is $O(1/\sqrt{N})$ in general, and $\exp(-\Omega(N))$ under a local-stability condition. Our proof is based on a framework from dynamic control known as \emph{dissipativity}. Our solution easy to implement and performs very well in practice when compared to the state of the art. Further, both our solution and our proof methodology can easily be generalized to more general constrained MDP settings and should thus, be of great interest to the burgeoning RMAB community.
On Dissipativity of Cross-Entropy Loss in Training ResNets
Püttschneider, Jens, Faulwasser, Timm
For example, the backpropagation in neural network training appears in optimal control as adjoint sensitivity equation (Esteve-Yagüe and Geshkovski, 2023; Faulwasser et al., 2021; Esteve et al., 2021). Moreover, the training of Neural Networks (NNs) with constant width in each layer can be formulated as an Optimal Control Problem (OCP) (Li et al., 2018; Esteve et al., 2021). In this context, the layer-to-layer propagation of the data is considered a dynamical system on a finite horizon corresponding to the depth of the network. The solution to the OCP determines the weights and biases of the neural network, i.e., weights and biases are the control inputs to drive the data to a desired point in the terminal layer determined by the label and loss function. In particular, the system and control perspective is helpful for Residual Neural Networks (ResNets) (He et al., 2016), which can be regarded as Euler forward discretizations of neural ODEs (Chen et al., 2018). Chang et al., 2018 analyze the reversibility and stability of the ResNet dynamics based on their continuous time counterparts. Esteve et al., 2021 and Faulwasser et al., 2021 have suggested to include a regularization term based on the states of the hidden layers in the training OCP. Faulwasser et al., 2021 analyze the dissipativity and the related turnpike property from an optimal control point of view when utilizing a quadratic (l
Synthesizing Neural Network Controllers with Closed-Loop Dissipativity Guarantees
Junnarkar, Neelay, Arcak, Murat, Seiler, Peter
In this paper, a method is presented to synthesize neural network controllers such that the feedback system of plant and controller is dissipative, certifying performance requirements such as L2 gain bounds. The class of plants considered is that of linear time-invariant (LTI) systems interconnected with an uncertainty, including nonlinearities treated as an uncertainty for convenience of analysis. The uncertainty of the plant and the nonlinearities of the neural network are both described using integral quadratic constraints (IQCs). First, a dissipativity condition is derived for uncertain LTI systems. Second, this condition is used to construct a linear matrix inequality (LMI) which can be used to synthesize neural network controllers. Finally, this convex condition is used in a projection-based training method to synthesize neural network controllers with dissipativity guarantees. Numerical examples on an inverted pendulum and a flexible rod on a cart are provided to demonstrate the effectiveness of this approach.
Unconstrained Parametrization of Dissipative and Contracting Neural Ordinary Differential Equations
Martinelli, Daniele, Galimberti, Clara Lucía, Manchester, Ian R., Furieri, Luca, Ferrari-Trecate, Giancarlo
In this work, we introduce and study a class of Deep Neural Networks (DNNs) in continuous-time. The proposed architecture stems from the combination of Neural Ordinary Differential Equations (Neural ODEs) with the model structure of recently introduced Recurrent Equilibrium Networks (RENs). We show how to endow our proposed NodeRENs with contractivity and dissipativity -- crucial properties for robust learning and control. Most importantly, as for RENs, we derive parametrizations of contractive and dissipative NodeRENs which are unconstrained, hence enabling their learning for a large number of parameters. We validate the properties of NodeRENs, including the possibility of handling irregularly sampled data, in a case study in nonlinear system identification.