e034fb6b66aacc1d48f445ddfb08da98-Reviews.html

Neural Information Processing Systems 

This paper presents a new method for using human feedback to improve a reinforcement-learning agent. The novelty of the approach is to transform human feedback into a potentially inconsistent estimate on the optimality of an action, instead of a reward as is often the case. The resultant algorithm outperforms the previous state of the art in a pair of toy domains. I thought this was an excellent paper, which appropriately motivated the problem, clearly introduced a new idea and then compared performance to other state-of-the-art algorithms (and not just strawmen). I mostly have suggestions for improvement.