Reinforcement Learning
Adversarial Attacks on Neural Network Policies
Huang, Sandy, Papernot, Nicolas, Goodfellow, Ian, Duan, Yan, Abbeel, Pieter
Machine learning classifiers are known to be vulnerable to inputs maliciously constructed by adversaries to force misclassification. Such adversarial examples have been extensively studied in the context of computer vision applications. In this work, we show adversarial attacks are also effective when targeting neural network policies in reinforcement learning. Specifically, we show existing adversarial example crafting techniques can be used to significantly degrade test-time performance of trained policies. Our threat model considers adversaries capable of introducing small perturbations to the raw input of the policy. We characterize the degree of vulnerability across tasks and training algorithms, for a subclass of adversarial-example attacks in white-box and black-box settings. Regardless of the learned task or training algorithm, we observe a significant drop in performance, even with small adversarial perturbations that do not interfere with human perception. Videos are available at http://rll.berkeley.edu/adversarial.
Multi-Focus Attention Network for Efficient Deep Reinforcement Learning
Choi, Jinyoung (Seoul National University) | Lee, Beom-Jin (Seoul National University) | Zhang, Byoung-Tak (Seoul National University)
Deep reinforcement learning (DRL) has shown incredible performance in learning various tasks to the human level. However, unlike human perception, current DRL models connect the entire low-level sensory input to the state-action values rather than exploiting the relationship between and among entities that constitute the sensory input. Because of this difference, DRL needs vast amount of experience samples to learn. In this paper, we propose a Multi-focus Attention Network (MANet) which mimics human ability to spatially abstract the low-level sensory input into multiple entities and attend to them simultaneously. The proposed method first divides the low-level input into several segments which we refer to as partial states. After this segmentation, parallel attention layers attend to the partial states relevant to solving the task. Our model estimates state-action values using these attended partial states. In our experiments, MANet attains highest scores with significantly less experience samples. Additionally, the model shows higher performance compared to the Deep Q-network and the single attention model as benchmarks. Furthermore, we extend our model to attentive communication model for performing multi-agent cooperative tasks. In multi-agent cooperative task experiments, our model shows 20% faster learning than existing state-of-the-art model.
Using Options to Accelerate Learning of New Tasks According to Human Preferences
Bonini, Rodrigo Cesar (Universidade de Sao Paulo) | Silva, Felipe Leno da (Universidade de Sรฃo Paulo) | Spina, Edison (Universidade de Sรฃo Paulo) | Costa, Anna Helena Reali (Universidade de Sรฃo Paulo)
Over the years, people need to incorporate a wider range of information and multiple objectives for their decision making. Nowadays, humans are dependent on computer systems to interpret and take profit from the huge amount of available data on the Internet. Hence, varied services, such as location-based systems, must combine a huge quantity of raw data to give the desired response to the user. However, as humans have different preferences, the optimal answer is different for each user profile, and few systems offer the service of solving tasks in a customized manner for each user. Reinforcement Learning (RL) has been used to autonomously train systems to solve (or assist on) decision-making tasks according to user preferences. However, the learning process is very slow and require many interactions with the environment. Therefore, we here propose to reuse knowledge from previous tasks to accelerate the learning process in a new task. Our proposal, called Multiobjective Options, accelerates learning while providing a customized solution according to the current user preferences. Our experiments in the Tourist World Domain show that our proposal learns faster and better than regular learning, and that the achieved solutions follow user preferences.
Causal Learning versus Reinforcement Learning for Knowledge Learning and Problem Solving
Ho, Seng-Beng (Institute of High Performance Computing)
Causal learning and reinforcement learning are both important AI learning mechanisms but are usually treated separately, despite the fact that both are directly relevant to problem solving processes. In this paper we propose a method for causal learning and problem solving, and compare and contrast that with AI reinforcement learning and show that the two methods are actually related, differing only in the values of the learning rate ฮฑ and discount factor ฮณ. However, the causal learning framework emphasizes quick but non-optimal concoction of problem solutions while AI reinforcement learning generates optimal solutions at the expense of speed. Cognitive science literature is reviewed and it is found that psychological reinforcement learning in lower form animals such as mammals is distinct from AI reinforcement learning in that psychological reinforcement learning strives neither for speed nor optimality, and that higher form animals such as humans and primates employ quick causal learning for survival instead of reinforcement learning. AI systems should likewise take advantage of a framework that employs rapid inductive causal learning to generate problem solutions for its general viability in terms of rapid adaptability, without the need to always strive for optimality.
Clyde: A Deep Reinforcement Learning DOOM Playing Agent
Ratcliffe, Dino Stephen (University of Essex) | Devlin, Sam (University of York) | Kruschwitz, Udo (University of Essex) | Citi, Luca (University of Essex)
In many cases games provide noise free computer science at Poznan University. It provides an interface environments and can also encompass the whole world state for AI agents to learn from the raw visual data that is in data structures easily. Much of the early work in this produced by DOOM (Kempka et al. 2016). They also run a domain has focussed on digital implementations of board competition that places these agents into death matches in games, such as backgammon (Tesauro 1995), chess (Campbell, order to compare their performance. A death match in the Hoane, and Hsu 2002) and more recently go (Silver case of this competition is a time limited game mode where et al. 2016). These games have then been used to benchmark each agent must accumulate the highest score possible by many different approaches, including tree search approaches killing other agents in the match. This is where our agent was such as Monte Carlo Tree Search (MCTS) (Browne et al. submitted in order to assess its performance against other 2012) along with other approaches such as deep reinforcement agents.
Value Alignment or Misalignment -- What Will Keep Systems Accountable?
Arnold, Thomas (Tufts University) | Kasenberg, Daniel (Tufts University) | Scheutz, Matthias (Tufts University)
Machine learning's advances have led to new ideas about the feasibility and importance of machine ethics keeping pace, with increasing emphasis on safety, containment, and alignment. This paper addresses a recent suggestion that inverse reinforcement learning (IRL) could be a means to so-called "value alignment.'' We critically consider how such an approach can engage the social, norm-infused nature of ethical action and outline several features of ethical appraisal that go beyond simple models of behavior, including unavoidably temporal dimensions of norms and counterfactuals. ย We propose that a hybrid approach for computational architectures still offers the most promising avenue for machines acting in an ethical fashion.
Free Data Science eBooks - February 2017
Reinforcement learning, one of the most active research areas in artificial intelligence, is a computational approach to learning whereby an agent tries to maximize the total amount of reward it receives when interacting with a complex, uncertain environment. In Reinforcement Learning, Richard Sutton and Andrew Barto provide a clear and simple account of the key ideas and algorithms of reinforcement learning. The only necessary mathematical background is familiarity with elementary concepts of probability. The book is divided into three parts. Part I defines the reinforcement learning problem in terms of Markov decision processes.
Artificial Intelligence's Next Big Step: Reinforcement Learning - The New Stack
Almost every machine learning breakthrough you hear about (and most of what's currently called "artificial intelligence") is supervised learning; where you start with a curated and labeled data set. But another technique, reinforcement learning, is just starting to make its way out of the research lab. Reinforcement learning is where an agent learns by interacting with its environment. It isn't told by a trainer what to do and it learns what actions to take to get the highest reward in the situation by trial and error, even when the reward isn't obvious and immediate. It learns how to solve problems rather than being taught what solutions look like. Reinforcement learning is how DeepMind created the AlphaGo system that beat a high-ranking Go player (and has recently been winning online Go matches anonymously). It's how University of California Berkeley's BRETT robot learns how to move its hands and arms to perform physical tasks like stacking blocks or screwing the lid onto a bottle, in just three hours (or ten minutes if it's told where the objects are that it's going to work with, and where they need to end up).
Learning Policies For Learning Policies -- Meta Reinforcement Learning (RLยฒ) in Tensorflow
Reinforcement Learning provides a framework for training agents to solve problems in the world. One of the limitations of these agents however is their inflexibility once trained. They are able to learn a policy to solve a specific problem (formalized as an MDP), but that learned policy is often useless in new problems, even relatively similar ones. Imagine the simplest possible agent: one trained to solve a two-armed bandit task in which one arm always provides a positive reward, and the other arm always provides no reward. Using any RL algorithm such as Q-Learning or Policy Gradient, the agent can quickly learn to always choose the arm with the positive reward.
Artificial Intelligence: Reinforcement Learning in Python
When people talk about artificial intelligence, they usually don't mean supervised and unsupervised machine learning. These tasks are pretty trivial compared to what we think of AIs doing - playing chess and Go, driving cars, and beating video games at a superhuman level. Reinforcement learning has recently become popular for doing all of that and more. Much like deep learning, a lot of the theory was discovered in the 70s and 80s but it hasn't been until recently that we've been able to observe first hand the amazing results that are possible. In 2016 we saw Google's AlphaGo beat the world Champion in Go. We saw AIs playing video games like Doom and Super Mario.