Addressing Function Approximation Error in Actor-Critic Methods

Fujimoto, Scott, van Hoof, Herke, Meger, Dave

Feb-26-2018–arXiv.org Machine Learning

In value-based reinforcement learning methods such as deep Q-learning, function approximation errors are known to lead to overestimated value estimates and suboptimal policies. We show that this problem persists in an actor-critic setting and propose novel mechanisms to minimize its effects on both the actor and critic. Our algorithm takes the minimum value between a pair of critics to restrict overestimation and delays policy updates to reduce per-update error. We evaluate our method on the suite of OpenAI gym tasks, outperforming the state of the art in every environment tested.

artificial intelligence, machine learning, reinforcement learning, (14 more...)

arXiv.org Machine Learning

Feb-26-2018

arXiv.org PDF

Add feedback

Country:
- North America > Canada (0.28)

Genre:
- Research Report (1.00)

Technology:
- Information Technology > Artificial Intelligence
  - Machine Learning > Reinforcement Learning (1.00)
  - Representation & Reasoning > Uncertainty
    - Fuzzy Logic (0.63)

Duplicate Docs Excel Report

Title
None found

Similar Docs Excel Report more

Title	Similarity	Source
None found