Goto

Collaborating Authors

 Learning Graphical Models




OnEfficiencyinHierarchicalReinforcement Learning

Neural Information Processing Systems

While this has been demonstrated empirically overtimeinavarietyoftasks,theoretical resultsquantifying thebenefits of such methods are still few and far between. In this paper, we discuss the kind of structure in a Markov decision process which gives rise to efficient HRLmethods.








LobsDICE: OfflineLearningfromObservationvia StationaryDistributionCorrectionEstimation

Neural Information Processing Systems

We additionally assume that the agent cannot interact with the environment but has access to the action-labeled transition data collected by some agents with unknown qualities.