Reviews: The Option Keyboard: Combining Skills in Reinforcement Learning

Neural Information Processing Systems 

Post Response update: Thank you for the detailed response. I still believe that a more in depth discussion of the differences or similarities of policy and cumulant based formulations is required to place the paper appropriately in context of prior work. I think the new results presented by the authors in the response partially address my concerns about comparisons with prior work but not fully. I would still like to see comparison against a policy-based method as per the authors' classification. I agree that all methods might have negative transfer but it would be ideal to include a discussion of the conditions under which the methods would show positive or negative transfer (something that the authors do) and to place that in context with other methods at least qualitatively (something that the authors dont). The newer evaluations in the response do satisfy a part of my concerns.