Continuous MDP Homomorphisms and Homomorphic Policy Gradient

Jan-15-2025, 16:36:29 GMT–Neural Information Processing Systems

Abstraction has been widely studied as a way to improve the efficiency and generalization of reinforcement learning algorithms. In this paper, we study abstraction in the continuous-control setting. We extend the definition of MDP homomorphisms to encompass continuous actions in continuous state spaces. We derive a policy gradient theorem on the abstract MDP, which allows us to leverage approximate symmetries of the environment for policy optimization. Based on this theorem, we propose an actor-critic algorithm that is able to learn the policy and the MDP homomorphism map simultaneously, using the lax bisimulation metric.

algorithm, continuous mdp homomorphism, homomorphism and homomorphic policy gradient, (1 more...)

Neural Information Processing Systems

Jan-15-2025, 16:36:29 GMT

Conferences Web Page

Add feedback

Technology:
- Information Technology > Artificial Intelligence > Machine Learning > Reinforcement Learning (0.66)