RUDDER: Return Decomposition for Delayed Rewards

Jose A. Arjona-Medina, Michael Gillhofer, Michael Widrich, Thomas Unterthiner, Johannes Brandstetter, Sepp Hochreiter

Mar-22-2025, 19:30:29 GMT–Neural Information Processing Systems

We propose RUDDER, a novel reinforcement learning approach for delayed rewards in finite Markov decision processes (MDPs). In MDPs the Q-values are equal to the expected immediate reward plus the expected future rewards. The latter are related to bias problems in temporal difference (TD) learning and to high variance problems in Monte Carlo (MC) learning. Both problems are even more severe when rewards are delayed. RUDDER aims at making the expected future rewards zero, which simplifies Q-value estimation to computing the mean of the immediate reward. We propose the following two new concepts to push the expected future rewards toward zero.

artificial intelligence, machine learning, reinforcement learning, (16 more...)

Neural Information Processing Systems

Mar-22-2025, 19:30:29 GMT

Conferences PDF

Add feedback

Country:
- North America > United States (0.46)

Industry:
- Leisure & Entertainment (0.30)

Technology:
- Information Technology > Artificial Intelligence > Machine Learning
  - Learning Graphical Models > Undirected Networks
    - Markov Models (0.34)
  - Neural Networks > Deep Learning (1.00)
  - Reinforcement Learning (1.00)

Duplicate Docs Excel Report

Title
RUDDER: Return Decomposition for Delayed Rewards

Similar Docs Excel Report more

Title	Similarity	Source
None found