Goto

Collaborating Authors

 Reinforcement Learning



Teaching Inverse Reinforcement Learners via Features and Demonstrations

Neural Information Processing Systems

Weintroduceanaturalquantity,the teaching risk, which measures the potential suboptimality of policies that look optimal to the learner in this setting. We show that bounds on the teaching risk guarantee that the learner is able to find a near-optimal policy using standard algorithms basedoninversereinforcement learning. Basedonthesefindings, we suggest a teaching scheme in which the expert can decrease the teaching risk by updating the learner's worldview, and thus ultimately enable her to find a near-optimalpolicy.



eda9523faa5e7191aee1c2eaff669716-Supplemental-Conference.pdf

Neural Information Processing Systems

Though promising results have been reported on some RL application domains, policies learned with such representations usually fail to generalize well in a complex environment because minimizing a reconstruction loss may potentially introduce local (visual) features with task-irrelevant information.


eda9523faa5e7191aee1c2eaff669716-Paper-Conference.pdf

Neural Information Processing Systems

Though promising results have been reported on some RL application domains, policies learned with such representations usually fail to generalize well in a complex environment because minimizing a reconstruction loss may potentially introduce local (visual) features with task-irrelevant information.



Object-CategoryAwareReinforcementLearning

Neural Information Processing Systems

Reinforcement Learning (RL) has achievedimpressiveprogress inrecent years, such asresults in Atari [24] and Go [28] in which RL agents even perform better than human beings.




A Novel Framework for Policy Mirror Descent with General Parameterization and Linear Convergence Carlo Alfano Department of Statistics University of Oxford

Neural Information Processing Systems

In this work, we introduce a framework for policy optimization based on mirror descent that naturally accommodates general parameterizations. The policy class induced by our scheme recovers known classes, e.g., softmax, and generates new ones depending on the choice of mirror map.