State-Regularized Policy Search for Linearized Dynamical Systems

Abdulsamad, Hany (Technische Universität Darmstadt) | Arenz, Oleg (Technische Universität Darmstadt) | Peters, Jan (Technische Universität Darmstadt) | Neumann, Gerhard (University of Lincoln)

Jun-14-2017–AAAI Conferences

Stability of the policy update is a major issue for these methods, rendering them hard to apply for highly nonlinear systems. Recent approaches combine classical Stochastic Optimal Control methods with information-theoretic bounds to control the step-size of the policy update and could even be used to train nonlinear deep control policies. These methods bound the relative entropy between the new and the old policy to ensure a stable policy update. However, despite the bound in policy space, the state distributions of two consecutive policies can still differ significantly, rendering the used local approximate models invalid. To alleviate this issue we propose enforcing a relative entropy constraint not only on the policy update, but also on the update of the state distribution, around which the dynamics and cost are being approximated. We present a derivation of the closed-form policy update and show that our approach outperforms related methods on two nonlinear and highly dynamic simulated systems.

constraint, policy update, state distribution, (15 more...)

AAAI Conferences

Jun-14-2017

Conferences PDF

Add feedback

Country:
- Europe
  - United Kingdom > England
    - Lincolnshire > Lincoln (0.04)
  - Germany > Hesse
    - Darmstadt Region > Darmstadt (0.05)

Technology:
- Information Technology > Artificial Intelligence
  - Machine Learning (1.00)
  - Representation & Reasoning > Optimization (0.95)
  - Robots (0.90)

Duplicate Docs Excel Report

Title
None found

Similar Docs Excel Report more

Title	Similarity	Source
None found