Risk-Sensitive Reinforcement Learning with Exponential Criteria

Noorani, Erfaun, Mavridis, Christos, Baras, John

Dec-19-2023–arXiv.org Artificial Intelligence

While reinforcement learning has shown experimental success in a number of applications, it is known to be sensitive to noise and perturbations in the parameters of the system, leading to high variance in the total reward amongst different episodes on slightly different environments. To introduce robustness, as well as sample efficiency, risk-sensitive reinforcement learning methods are being thoroughly studied. In this work, we provide a definition of robust reinforcement learning policies and formulate a risk-sensitive reinforcement learning problem to approximate them, by solving an optimization problem with respect to a modified objective based on exponential criteria. In particular, we study a model-free risksensitive variation of the widely-used Monte Carlo Policy Gradient algorithm, and introduce a novel risk-sensitive online Actor-Critic algorithm based on solving a multiplicative Bellman equation using stochastic approximation updates. Analytical results suggest that the use of exponential criteria generalizes commonly used ad-hoc regularization approaches, improves sample efficiency, and introduces robustness with respect to perturbations in the model parameters and the environment. The implementation, performance, and robustness properties of the proposed methods are evaluated in simulated experiments.

algorithm, reinforcement learning, risk-sensitive reinforcement learning, (10 more...)

arXiv.org Artificial Intelligence

Dec-19-2023

arXiv.org PDF

Add feedback

Country:
- North America > United States
  - Maryland > Prince George's County
    - College Park (0.04)
  - Louisiana > Orleans Parish
    - New Orleans (0.04)
  - California > Los Angeles County
    - Santa Monica (0.04)
- Europe
  - United Kingdom > Scotland
    - City of Edinburgh > Edinburgh (0.04)
  - Sweden > Stockholm
    - Stockholm (0.04)
  - France > Hauts-de-France
    - Nord > Lille (0.04)
- Asia > Middle East
  - Jordan (0.04)

Genre:
- Research Report > New Finding (0.34)

Industry:
- Education (0.66)

Technology:
- Information Technology > Artificial Intelligence > Machine Learning
  - Reinforcement Learning (1.00)
  - Statistical Learning > Gradient Descent (0.46)