
While this has been demonstrated empirically overtimeinavarietyoftasks,theoretical resultsquantifying thebenefits of such methods are still few and far between. In this paper, we discuss the kind of structure in a Markov decision process which gives rise to efficient HRLmethods.