Quantile Constrained Reinforcement Learning: A Reinforcement Learning Framework Constraining Outage Probability

Oct-10-2024, 10:56:58 GMT–Neural Information Processing Systems

Constrained reinforcement learning (RL) is an area of RL whose objective is to find an optimal policy that maximizes expected cumulative return while satisfying a given constraint. Most of the previous constrained RL works consider expected cumulative sum cost as the constraint. However, optimization with this constraint cannot guarantee a target probability of outage event that the cumulative sum cost exceeds a given threshold. This paper proposes a framework, named Quantile Constrained RL (QCRL), to constrain the quantile of the distribution of the cumulative sum cost that is a necessary and sufficient condition to satisfy the outage constraint. This is the first work that tackles the issue of applying the policy gradient theorem to the quantile and provides theoretical results for approximating the gradient of the quantile.

cumulative sum cost, learning framework constraining outage probability, quantile constrained reinforcement learning, (2 more...)

Neural Information Processing Systems

Oct-10-2024, 10:56:58 GMT

Conferences Web Page

Add feedback

Technology:
- Information Technology > Artificial Intelligence > Machine Learning > Reinforcement Learning (1.00)