Reinforcement Learning
On- and Off-Policy Monotonic Policy Improvement
Monotonic policy improvement and off-policy learning are two main desirable properties for reinforcement learning algorithms. In this paper, by lower bounding the performance difference of two policies, we show that the monotonic policy improvement is guaranteed from on- and off-policy mixture samples. An optimization procedure which applies the proposed bound can be regarded as an off-policy natural policy gradient method. In order to support the theoretical result, we provide a trust region policy optimization method using experience replay as a naive application of our bound, and evaluate its performance in two classical benchmark problems.
Learning Hard Alignments with Variational Inference
Lawson, Dieterich, Chiu, Chung-Cheng, Tucker, George, Raffel, Colin, Swersky, Kevin, Jaitly, Navdeep
There has recently been significant interest in hard attention models for tasks such as object recognition, visual captioning and speech recognition. Hard attention can offer benefits over soft attention such as decreased computational cost, but training hard attention models can be difficult because of the discrete latent variables they introduce. Previous work used REINFORCE and Q-learning to approach these issues, but those methods can provide high-variance gradient estimates and be slow to train. In this paper, we tackle the problem of learning hard attention for a sequential task using variational inference methods, specifically the recently introduced VIMCO and NVIL. Furthermore, we propose a novel baseline that adapts VIMCO to this setting. We demonstrate our method on a phoneme recognition task in clean and noisy environments and show that our method outperforms REINFORCE, with the difference being greater for a more complicated task.
Integrating Knowledge Representation, Reasoning, and Learning for Human-Robot Interaction
Sridharan, Mohan (The University of Auckland)
Robots interacting with humans often have to represent and reason with different descriptions of incomplete domain knowledge and uncertainty, and revise this knowledge over time. Towards achieving these capabilities, the architecture described in this paper combines the complementary strengths of declarative programming, probabilistic graphical models, and reinforcement learning. For any given goal, non-monotonic logical reasoning with a coarse-resolution representation of the domain is used to compute a tentative plan of abstract actions. Each abstract action is implemented as a sequence of concrete actions by reasoning probabilistically over the relevant part of a fine-resolution representation tightly-coupled to the coarse-resolution representation. The outcomes of executing the concrete actions are used for subsequenct reasoning at the coarse resolution. Furthermore, the task of interactively learning axioms governing action capabilities, preconditions and effects, is posed as a relational reinforcement learning problem, using decision tree regression and sampling to construct and generalize over candidate axioms. These capabilities are illustrated in simulation and on a physical robot moving objects to specific people or locations in an indoor domain.
An Integrated Computational Framework for Attention, Reinforcement Learning, and Working Memory
Stocco, Andrea (University of Washington)
This paper proposes a reinterpretation of selective attention as a form of control of working memory based on self-generated reward signals and model-free reinforcement learning. In addition to being simple and parsimonious, this approach systematizes a number of classic psychological constructs without calling for additional, specific mechanisms. Finally, the papers presents the results of an empirical test of this framework, and elaborates on the implications of our findings for general models of control and intelligent behavior, as well as neurobiological models of the basal ganglia.
A Framework Using Machine Vision and Deep Reinforcement Learning for Self-Learning Moving Objects in a Virtual Environment
Wu, Richard (University of Massachusetts Dartmouth) | Zhao, Ying (Naval Postgraduate School) | Clarke, Alan (Naval Postgraduate School) | Kendall, Anthony (Naval Postgraduate School)
In recent artificial intelligence (AI) research, convolutional neural networks (CNNs) can create artificial agents capable of self-learning. Self-learning autonomous moving objects utilize machine vision techniques based on processing and recognizing objects in digital images. Afterwards, deep reinforcement learning (Deep-RL) is applied to understand and learn intelligent actions and controls. The objective of our research is to study methods and designs on how machine vision and deep machine learning algorithms can be implemented in a virtual world (e.g., a computer game) for moving objects (e.g., vehicles or aircrafts) to improve their navigation and detection of threats in real life. In this paper, we create a framework for generating and using data from computer games to be used in CNNs and Deep-RL to perform intelligent actions. We show the initial results of applying the framework and identify various military applications that may benefit from this research.
Toward Supervised Reinforcement Learning with Partial States for Social HRI
Senft, Emmanuel (Plymouth University) | Lemaignan, Sรฉverin (Plymouth University) | Baxter, Paul (University of Lincoln) | Belpaeme, Tony (Plymouth University)
Social interacting is a complex task for which machine learning holds particular promise. However, as no sufficiently accurate simulator of human interactions exists today, the learning of social interaction strategies has to happen online in the real world. Actions executed by the robot impact on humans, and as such have to be carefully selected, making it impossible to rely on random exploration. Additionally, no clear reward function exists for social interactions. This implies that traditional approaches used for Reinforcement Learning cannot be directly applied for learning how to interact with the social world. As such we argue that robots will profit from human expertise and guidance to learn social interactions. However, as the quantity of input a human can provide is limited, new methods have to be designed to use human input more efficiently. In this paper we describe a setup in which we combine a framework called Supervised Progressively Autonomous Robot Competencies (SPARC), which allows safer online learning with Reinforcement Learning, with the use of partial states rather than full states to accelerate generalisation and obtain a usable action policy more quickly.
Toward Probabilistic Safety Bounds for Robot Learning from Demonstration
Brown, Daniel S. (University of Texas at Austin) | Niekum, Scott (University of Texas at Austin)
Learning from demonstration is a popular method for teaching robots new skills. However, little work has looked at how to measure safety in the context of learning from demonstrations. We discuss three different types of safety problems that are important for robot learning from human demonstrations: (1) using demonstrations to evaluate the safety of a robot's current policy, (2) using demonstrations to enable risk-aware policy improvement, and (3) determining when the demonstrations received by the robot are sufficient to ensure a desired safety level. We propose a risk-aware Bayesian sampling approach based on inverse reinforcement learning that provides a first step towards addressing these problems. We demonstrate the validity of our approach on a simulated navigation task and discuss promising areas for future work.
A Goal-Based Movement Model for Continuous Multi-Agent Tasks
Despite increasing attention paid to the need for fast, scalable methods to analyze next-generation neuroscience data, comparatively little attention has been paid to the development of similar methods for behavioral analysis. Just as the volume and complexity of brain data have grown, behavioral paradigms in systems neuroscience have likewise become more naturalistic and less constrained, necessitating an increase in the flexibility and scalability of the models used to study them. In particular, key assumptions made in the analysis of typical decision paradigms --- optimality; analytic tractability; discrete, low-dimensional action spaces --- may be untenable in richer tasks. Here, using the case of a two-player, real-time, continuous strategic game as an example, we show how the use of modern machine learning methods allows us to relax each of these assumptions. Following an inverse reinforcement learning approach, we are able to succinctly characterize the joint distribution over players' actions via a generative model that allows us to simulate realistic game play. We compare simulated play from a number of generative time series models and show that ours successfully resists mode collapse while generating trajectories with the rich variability of real behavior. Together, these methods offer a rich class of models for the analysis of continuous action tasks at the single-trial level.
Projective simulation with generalization
Melnikov, Alexey A., Makmal, Adi, Dunjko, Vedran, Briegel, Hans J.
The ability to act upon a new stimulus, based on previous experience with similar, but distinct, stimuli, sometimes denoted as generalization, is used extensively in our daily life. As a simple example, consider a driver's response to traffic lights: The driver need not recognize the details of a particular traffic light in order to respond to it correctly, even though traffic lights may appear different from one another. The only property that matters is the color, whereas neither shape nor size should play any role in the driver's reaction. Learning how to react to traffic lights thus involves an aspect of generalization. A learning agent, capable of a meaningful and useful generalization is expected to have the following characteristics: (a) an ability for categorization (recognizing that all red signals have a common property, which we can refer to as redness); (b) an ability to classify (a new red object is to be related to the group of objects with the redness property); (c) ideally, only generalizations that are relevant for the success of the agent should be learned (red signals should be treated the same, whereas squareshaped signals should not, as they share no property that is of relevance in this context); (d) correct actions should be associated with relevant generalized properties (the driver should stop whenever a red signal is shown); and (e) the generalization mechanism should be flexible. To illustrate what we mean by "flexible generalization", let us go back to our driver. After learning how to handle traffic lights correctly, the driver tries to follow arrow signs to, say, a nearby airport. Clearly, it is now the shape category of the signal that should guide the driver, rather than the color category.