Q-learning as a monotone scheme
–arXiv.org Artificial Intelligence
Stability issues with reinforcement learning methods persist. To better understand some of these stability and convergence issues involving deep reinforcement learning methods, we examine a simple linear quadratic example. We interpret the convergence criterion of exact Q-learning in the sense of a monotone scheme and discuss consequences of function approximation on monotonicity properties.
arXiv.org Artificial Intelligence
May-30-2024
- Country:
- Europe > United Kingdom > England > Oxfordshire > Oxford (0.04)
- Genre:
- Research Report (0.50)
- Technology: