On Convergence and Rate of Convergence of Policy Improvement Algorithms
Ma, Jin, Wang, Gaozhan, Zhang, Jianfeng
–arXiv.org Artificial Intelligence
The Policy Improvement Algorithm (PIA) is a well-known approach in numerical optimal control theory, see, e.g., Jacka-Mijatović [5], Kerimkulov-Siska-Szpruch [6, 7], Puterman [10]. Its main idea is to construct an iteration scheme for the control actions that traces the maximizers/minimizers of the Hamiltonian, so that the corresponding returns are naturally improving. Mathematically this amounts to a type of Picard iteration for the associated HJB equations. Motivated by the above problem but with model uncertainty, the Reinforcement Learning (RL) algorithms for the entropy-regularized stochastic control problems have received very strong attention in recent years. By using relaxed control, the control problem is regularized or say penalized by Shannon's entropy, which captures the trade-off between exploitation (to optimize) and exploration (to learn the model).
arXiv.org Artificial Intelligence
Jun-20-2024
- Country:
- North America > United States > California > Los Angeles County > Los Angeles (0.14)
- Genre:
- Research Report (0.40)
- Technology: