Conservative Offline Distributional Reinforcement Learning

Dec-24-2025, 14:52:49 GMT–Neural Information Processing Systems

Many reinforcement learning (RL) problems in practice are offline, learning purely from observational data. A key challenge is how to ensure the learned policy is safe, which requires quantifying the risk associated with different actions. In the online setting, distributional RL algorithms do so by learning the distribution over returns (i.e., cumulative rewards) instead of the expected return; beyond quantifying risk, they have also been shown to learn better representations for planning.

artificial intelligence, machine learning, reinforcement learning, (8 more...)

Neural Information Processing Systems

Dec-24-2025, 14:52:49 GMT

Conferences Web Page

Add feedback

Technology:
- Information Technology > Artificial Intelligence > Machine Learning > Reinforcement Learning (0.66)