Goto

Collaborating Authors

 Industry




Faster Deep Reinforcement Learning with Slower Online Network

Neural Information Processing Systems

Deep reinforcement learning algorithms often use two networks for value function optimization: an online network, and a target network that tracks the online network with some delay. Using two separate networks enables the agent to hedge against issues that arise when performing bootstrapping.



In most cases, the game designer is expected to first learn about the agents

Neural Information Processing Systems

We would like to thank all reviewers for reading our paper and providing constructive comments. Sometimes, the primary interest is to understand agent behaviors, and hence only the learning mode is needed. Alternatively, when all game inputs are known, the focus is on the intervention mode. In the final version, we will (i) explain in 2.1 how these We agree that it is neither rigorous nor necessary to assert that "most" Our work is inspired by the current interests on complex optimization-based layers. It is the first to treat VIs as individual layers in the end-to-end framework.