Goto

Collaborating Authors

 Markov Models



Average-Reward Learning and Planning with Options Yi Wan, Abhishek Naik, Richard S. Sutton {wan6,anaik1,rsutton }@ualberta.ca University of Alberta, Amii

Neural Information Processing Systems

We extend the options framework for temporal abstraction in reinforcement learning from discounted Markov decision processes (MDPs) to average-reward MDPs. Our contributions include general convergent off-policy inter-option learning algorithms, intra-option algorithms for learning values and models, as well as sample-based planning variants of our learning algorithms. Our algorithms and convergence proofs extend those recently developed by Wan, Naik, and Sutton.









We thank the reviewers for their time and thorough comments, as well as their valuation of our work including its

Neural Information Processing Systems

For the larger discussion items, please find the detailed comments below. Additionally, the reviewers highlighted the importance of quantitative fits. We currently attempt to differentiate between these models using additional manipulations. R-learning may be advantageous for computation. Our work builds upon results in the field including Ref [2] This observation enabled us to pursue the hypothesis of the leaky estimate of average reward.