Goto

Collaborating Authors

 Education



Incentivizing Combinatorial Bandit Exploration

Neural Information Processing Systems

Consider a bandit algorithm that recommends actions to self-interested users in a recommendation system. The users are free to choose other actions and need to be incentivized to follow the algorithm's recommendations.




Giving Feedback on Interactive Student Programs with Meta-Exploration

Neural Information Processing Systems

One approach toward automatic grading is to learn an agent that interacts with a student's program and explores states indicative of errors via reinforcement learning. However, existing work on this approach only provides binary feedback of whether a program is correct or not, while students require finer-grained feedback on the specific errors in their programs to understand their mistakes. In this work, we show that exploring to discover errors can be cast as a meta-exploration problem.