On Convergence and Rate of Convergence of Policy Improvement Algorithms

Ma, Jin, Wang, Gaozhan, Zhang, Jianfeng

arXiv.org Artificial Intelligence 

The Policy Improvement Algorithm (PIA) is a well-known approach in numerical optimal control theory, see, e.g., Jacka-Mijatović [5], Kerimkulov-Siska-Szpruch [6, 7], Puterman [10]. Its main idea is to construct an iteration scheme for the control actions that traces the maximizers/minimizers of the Hamiltonian, so that the corresponding returns are naturally improving. Mathematically this amounts to a type of Picard iteration for the associated HJB equations. Motivated by the above problem but with model uncertainty, the Reinforcement Learning (RL) algorithms for the entropy-regularized stochastic control problems have received very strong attention in recent years. By using relaxed control, the control problem is regularized or say penalized by Shannon's entropy, which captures the trade-off between exploitation (to optimize) and exploration (to learn the model).

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found