Goto

Collaborating Authors

 Reinforcement Learning


ProvablyGoodBatchReinforcementLearning WithoutGreatExploration

Neural Information Processing Systems

Thisisbecause, in the traditional analysis, the error bound scales up with this ratio. We show that using pessimistic value estimatesin the low-data regions in Bellman optimality and evaluation back-up can yield more adaptive and stronger guarantees when the concentrability assumption does not hold.



0ee633a6ade45eab4276352b3ee79c7a-Paper-Conference.pdf

Neural Information Processing Systems

A fundamental difference between our learning problem from standard RL problems is that the realized reward feedback from conversion incrementality ismixed and delayed.



AConsciousness-InspiredPlanningAgentfor Model-Based ReinforcementLearning

Neural Information Processing Systems

Whether when planning our paths home from the office or from a hotel to an airport in an unfamiliar city, we typically focus on a small subset of relevant variables,e.g. the changeinposition orthepresence oftraffic.


Conservative Q-Learning for Offline Reinforcement Learning A viral Kumar

Neural Information Processing Systems

Effectively leveraging large, previously collected datasets in reinforcement learning (RL) is a key challenge for large-scale real-world applications. Offline RL algorithms promise to learn effective policies from previously-collected, static datasets without further interaction.