Global Policy Construction in Modular Reinforcement Learning
Zhang, Ruohan (The University of Texas at Austin) | Song, Zhao (The University of Texas at Austin) | Ballard, Dana H. (The University of Texas at Austin)
We propose a modular reinforcement learning algorithm which decomposes a Markov decision process into independent modules. Each module is trained using Sarsa(lambda). We introduce three algorithms for forming global policy from modules policies, and demonstrate our results using a 2D grid world.
Mar-6-2015