Goto

Collaborating Authors

 Country



ProvablyGoodBatchReinforcementLearning WithoutGreatExploration

Neural Information Processing Systems

Thisisbecause, in the traditional analysis, the error bound scales up with this ratio. We show that using pessimistic value estimatesin the low-data regions in Bellman optimality and evaluation back-up can yield more adaptive and stronger guarantees when the concentrability assumption does not hold.






0ee633a6ade45eab4276352b3ee79c7a-Paper-Conference.pdf

Neural Information Processing Systems

A fundamental difference between our learning problem from standard RL problems is that the realized reward feedback from conversion incrementality ismixed and delayed.