Goto

Collaborating Authors

 Reinforcement Learning





COMBO: ConservativeOfflineModel-Based PolicyOptimization

Neural Information Processing Systems

Offline reinforcement learning (offline RL) [30,34]refers tothe setting where policies are trained using static, previously collected datasets. This presents an attractive paradigm for data reuse and safe policy learning in many applications, such as healthcare [62], autonomous driving [65], robotics [25, 48], and personalized recommendation systems [59].







Small batch deep reinforcement learning

Neural Information Processing Systems

Since the policy used to collect transitions is changing throughout learning, the replay memory contains data coming from a mixture of policies (that differ from the agent's current policy), and