PoPS: Policy Pruning and Shrinking for Deep Reinforcement Learning
–arXiv.org Artificial Intelligence
Abstract-- The recent success of deep neural networks (DNNs) for function approximation in reinforcement learning has t rig-gered the development of Deep Reinforcement Learning (DRL) algorithms in various fields, such as robotics, computer gam es, natural language processing, computer vision, sensing sys tems, and wireless networking. Unfortunately, DNNs suffer from h igh computational cost and memory consumption, which limits th e use of DRL algorithms in systems with limited hardware resources. In recent years, pruning algorithms have demonstrated cons id-erable success in reducing the redundancy of DNNs in classifi cation tasks. However, existing algorithms suffer from a sign ificant performance reduction in the DRL domain. In this paper, we develop the first effective solution to the performance redu ction problem of pruning in the DRL domain, and establish a working algorithm, named Policy Pruning and Shrinking (PoPS), to tr ain DRL models with strong performance while achieving a compac t representation of the DNN. The framework is based on a novel iterative policy pruning and shrinking method that leverag es the power of transfer learning when training the DRL model. We present an extensive experimental study that demonstrates the strong performance of PoPS using the popular Cartpole, Luna r Lander, Pong, and Pacman environments. Finally, we develop an open source software for the benefit of researchers and devel opers in related fields. Deep reinforcement learning (DRL) algorithms have attracted much attention in recent years due to their capabili ty to provide a good approximation of the objective value in decision making tasks while dealing with very large state an d action spaces. In contrast to classic reinforcement learni ng methods that perform well for small-size models but perform poorly for large-scale models, DRL combines a deep neural network (DNN) with reinforcement learning for overcoming this issue. The DNN is used to map from states to actions in large-scale models so as to yield a policy that maximizes the objective value. In DeepMind's recently published Natu re paper [1], [2], a DRL algorithm was developed to teach computers how to play Atari games directly from the on-scree n Personal use of this material is permitted. Dor Livne and Kobi Cohen are with the School of Electrical and Computer Engineering, Ben-Gurion University of the Negev, Beer Shev a 8410501 Israel. This work was supported in part by the U.S.-Israel Binationa l Science Foundation (BSF) under grant 2017723, and by the Cyber Secur ity Research Center at Ben-Gurion University of the Negev under grant 076 /16.
arXiv.org Artificial Intelligence
Jan-14-2020
- Country:
- Asia > Middle East > Israel (0.44)
- Genre:
- Research Report > New Finding (0.34)
- Industry:
- Education (1.00)
- Leisure & Entertainment > Games
- Computer Games (0.54)
- Technology: