Goto

Collaborating Authors

 Reinforcement Learning


e140dbab44e01e699491a59c9978b924-Paper.pdf

Neural Information Processing Systems

Success stories of deep reinforcement learning (RL) from high dimensional inputs such as pixels or large spatial layouts include achieving superhuman performance on Atari games [30, 37, 1], grandmaster levelinStarcraft II[50]andgrasping adiverse setofobjects with impressivesuccess rates and generalization with robots in the real world [21].









131f383b434fdf48079bff1e44e2d9a5-AuthorFeedback.pdf

Neural Information Processing Systems

See Table 1for the average running time per problem instance. Note that the implementation of Z3 and OR-tools22 are in C++, while NeuRewriter and RL baselines are in Python. Still, we can observethat our approach achieves a23 better balance between the time-efficiency and the result quality. For expression simplification and job scheduling,24 NeuRewriter is even more time-efficient than Z3 and OR-tools. The region-pickerπω is parameterized by aQ-function and is similar in spirit to soft-Q learning [2].