Reinforcement Learning
Cornhole: A Widely-Accessible AI Robotics Task
Derbinsky, Nate (Wentworth Institute of Technology) | Frasca, Tyler M. (Tufts University)
In this paper we present the game of cornhole as a compelling, accessible, and adaptable AI robotics task. Cornhole is a fun and social game with simple rules, but involves strategy and physical training for humans to play competitively; thus, developing a robot that can play at the level of even the average human player presents a multitude of opportunities for curricular integration at a variety of levels. We characterize the AI tasks involved with the game, and present results and resources gained from preliminary offerings.
Handwriting Profiling Using Generative Adversarial Networks
Ghosh, Arna (Indian Institute of Technology Kharagpur) | Bhattacharya, Biswarup (Indian Institute of Technology Kharagpur) | Chowdhury, Somnath Basu Roy (Indian Institute of Technology Kharagpur)
Handwriting is a skill learned by humans from a very early age. The ability to develop oneโs own unique handwriting as well as mimic another personโs handwriting is a task learned by the brain with practice. This paper deals with this very problem where an intelligent system tries to learn the handwriting of an entity using Generative Adversarial Networks (GANs). We propose a modified architecture of DCGAN (Radford, Metz, and Chintala 2015) to achieve this. We also discuss about applying reinforcement learning techniques to achieve faster learning. Our algorithm hopes to give new insights in this area and its uses include identification of forged documents, signature verification, computer generated art, digitization of documents among others. Our early implementation of the algorithm illustrates a good performance with MNIST datasets.
Arnold: An Autonomous Agent to Play FPS Games
Chaplot, Devendra Singh (Carnegie Mellon University) | Lample, Guillaume (Carnegie Mellon University)
Advances in deep reinforcement learning have allowed autonomous agents to perform well on Atari games, often outperforming humans, using only raw pixels to make their decisions. However, most of these games take place in 2D environments that are fully observable to the agent. In this paper, we present Arnold, a completely autonomous agent to play First-Person Shooter Games using only screen pixel data and demonstrate its effectiveness on Doom, a classical first-person shooter game. Arnold is trained with deep reinforcement learning using a recent Action-Navigation architecture, which uses separate deep neural networks for exploring the map and fighting enemies. Furthermore, it utilizes a lot of techniques such as augmenting high-level game features, reward shaping and sequential updates for efficient training and effective performance. Arnold outperforms average humans as well as in-built game bots on different variations of the deathmatch. It also obtained the highest kill-to-death ratio in both the tracks of the Visual Doom AI Competition and placed second in terms of the number of frags.
Learning Options in Multiobjective Reinforcement Learning
Bonini, Rodrigo Cesar (Escola Politรฉcnica da Universidade de Sรฃo Paulo) | Silva, Felipe Leno da (Escola Politรฉcnica da Universidade de Sรฃo Paulo) | Costa, Anna Helena Reali (Escola Politรฉcnica da Universidade de Sรฃo Paulo)
Reinforcement Learning (RL) is a successful technique to train autonomous agents. However, the classical RL methods take a long time to learn how to solve tasks. Option-based solutions can be used to accelerate learning and transfer learned behaviors across tasks by encapsulating a partial policy into an action. However, the literature report only single-agent and single-objective option-based methods, but many RL tasks, especially real-world problems, are better described through multiple objectives. We here propose a method to learn options in Multiobjective Reinforcement Learning domains in order to accelerate learning and reuse knowledge across tasks. Our initial experiments in the Goldmine Domain show that our proposal learn useful options that accelerate learning in multiobjective domains. Our next steps are to use the learned options to transfer knowledge across tasks and evaluate this method with stochastic policies.
Scalable Multitask Policy Gradient Reinforcement Learning
Bsat, Salam El (Rafik Hariri University) | Ammar, Haitham Bou (American University of Beirut) | Taylor, Matthew E. (Washington State University)
Policy search reinforcement learning (RL) allows agents to learn autonomously with limited feedback. However, such methods typically require extensive experience for successful behavior due to their tabula rasa nature. Multitask RL is an approach, which aims to reduce data requirements by allowing knowledge transfer between tasks. Although successful, current multitask learning methods suffer from scalability issues when considering large number of tasks. The main reasons behind this limitation is the reliance on centralized solutions. This paper proposes to a novel distributed multitask RL framework, improving the scalability across many different types of tasks. Our framework maps multitask RL to an instance of general consensus and develops an efficient decentralized solver. We justify the correctness of the algorithm both theoretically and empirically: we first proof an improvement of convergence speed to an order of O(1/k) with k being the number of iterations, and then show our algorithm surpassing others on multiple dynamical system benchmarks.
Where to Add Actions in Human-in-the-Loop Reinforcement Learning
Mandel, Travis (University of Washington) | Liu, Yun-En (Enlearn) | Brunskill, Emma (Carnegie Mellon University) | Popoviฤ, Zoran (University of Washington)
In order for reinforcement learning systems to learn quickly in vast action spaces such as the space of all possible pieces of text or the space of all images, leveraging human intuition and creativity is key. However, a human-designed action space is likely to be initially imperfect and limited; furthermore, humans may improve at creating useful actions with practice or new information. Therefore, we propose a framework in which a human adds actions to a reinforcement learning system over time to boost performance. In this setting, however, it is key that we use human effort as efficiently as possible, and one significant danger is that humans waste effort adding actions at places (states) that aren't very important. Therefore, we propose Expected Local Improvement (ELI), an automated method which selects states at which to query humans for a new action. We evaluate ELI on a variety of simulated domains adapted from the literature, including domains with over a million actions and domains where the simulated experts change over time. We find ELI demonstrates excellent empirical performance, even in settings where the synthetic "experts" are quite poor.
Learning to Act by Predicting the Future
Dosovitskiy, Alexey, Koltun, Vladlen
We present an approach to sensorimotor control in immersive environments. Our approach utilizes a high-dimensional sensory stream and a lower-dimensional measurement stream. The cotemporal structure of these streams provides a rich supervisory signal, which enables training a sensorimotor control model by interacting with the environment. The model is trained using supervised learning techniques, but without extraneous supervision. It learns to act based on raw sensory input from a complex three-dimensional environment. The presented formulation enables learning without a fixed goal at training time, and pursuing dynamically changing goals at test time. We conduct extensive experiments in three-dimensional simulations based on the classical first-person game Doom. The results demonstrate that the presented approach outperforms sophisticated prior formulations, particularly on challenging tasks. The results also show that trained models successfully generalize across environments and goals. A model trained using the presented approach won the Full Deathmatch track of the Visual Doom AI Competition, which was held in previously unseen environments.
This Week's Awesome Stories From Around the Web (Through February 11th)
Understanding Agent Cooperation Joel Leibo, Vinicius Zambaldi, Marc Lanctot, Janusz Marecki, Thore Graepel Google DeepMind Blog "Recent progress in artificial intelligence and specifically deep reinforcement learning provides us with the tools to look at the problem of social dilemmas through a new lens... we showed that we can apply the modern AI technique of deep multi-agent reinforcement learning to age-old questions in social science such as the mystery of the emergence of cooperation." Agility Robotics Introduces Cassie, a Dynamic and Talented Robot Delivery Ostrich Evan Ackerman IEEE Spectrum "Agility Robotics, a spin-off of Oregon State University, is officially announcing a shiny new bipedal robot named Cassie. Cassie is a dynamic walker, meaning that it walks much more like humans do than most of the carefully plodding bipedal robots we're used to seeing... Cassie has some work to do before it's ready to be hauling groceries up stairs for you, but we're very much looking forward to watching this robot taking more steps toward robust and dynamic legged locomotion." How Escape Rooms and Live Theater Are Paving the Way for VR Bryan Bishop The Verge "Cinema has had more than a century to develop its own language of shots, cuts, and transitions, while storytelling in VR is still in its infancy... creators seem to be zeroing in on interactive, experiential moments as one of the key building blocks of VR storytelling. One of Chris Milk's next projects is a piece set in the Planet of the Apes universe that will lean heavily on AI to drive interactive character performances."
Batch Policy Gradient Methods for Improving Neural Conversation Models
Kandasamy, Kirthevasan, Bachrach, Yoram, Tomioka, Ryota, Tarlow, Daniel, Carter, David
We study reinforcement learning of chatbots with recurrent neural network architectures when the rewards are noisy and expensive to obtain. For instance, a chatbot used in automated customer service support can be scored by quality assurance agents, but this process can be expensive, time consuming and noisy. Previous reinforcement learning work for natural language processing uses on-policy updates and/or is designed for on-line learning settings. We demonstrate empirically that such strategies are not appropriate for this setting and develop an off-policy batch policy gradient method (BPG). We demonstrate the efficacy of our method via a series of synthetic experiments and an Amazon Mechanical Turk experiment on a restaurant recommendations dataset.
Reinforcement Learning as a Service
I've been integrating reinforcement learning into an actual product for the last 6 months, and therefore I'm developing an appreciation for what are likely to be common problems. In particular, I'm now sold on the idea of reinforcement learning as a service, of which the decision service from MSR-NY is an early example (limited to contextual bandits at the moment, but incorporating key system insights). Service, not algorithm Supervised learning is essentially observational: some data has been collected and subsequently algorithms are run on it. In contrast, counterfactual learning is very difficult do to observationally. Diverse fields such as economics, political science, and epidemiology all attempt to make counterfactual conclusions using observational data, essentially because this is the only data available (at an affordable cost).