Education
Nonparametric Bayesian Learning of Other Agents' Policies in Interactive POMDPs
Panella, Alessandro (University of Illinois at Chicago) | Gmytrasiewicz, Piotr (University of Illinois at Chicago)
We consider an autonomous agent facing a partially observable, stochastic, multiagent environment where the unknown policies of other agents are represented as finite state controllers (FSCs). We show how an agent can (i) learn the FSCs of the other agents, and (ii) exploit these models during interactions. To separate the issues of off-line versus on-line learning we consider here an off-line two-phase approach. During the first phase the agent observes as the other player(s) are interacting with the environment (the observations may be imperfect and the learning agent is not taking part in the interaction.) The collected data is used to learn an ensemble of FSCs that explain the behavior of the other agent(s) using a Bayesian non-parametric (BNP) approach. We verify the quality of the learned models during the second phase by allowing the agent to compute its own optimal policy and interact with the observed agent. The optimal policy for the learning agent is obtained by solving an interactive POMDP in which the states are augmented by the other agent(s)' possible FSCs. The advantage of using the Bayesian nonparametric approach in the first phase is that the complexity (number of nodes) of the learned controllers is not bounded a priori. Our two-phase approach is preliminary and separates the learning using BNP from the complexities of learning on-line while the other agent may be modifying its policy (on-line approach is subject of our future work.) We describe our implementation and results in a multiagent Tiger domain. Our results show that learning improves the agent's performance, which increases with the amount of data collected during the learning phase.
What Women Want: Analyzing Research Publications to Understand Gender Preferences in Computer Science
Mihalcea, Rada (University of Michigan) | Welch, Charles (University of Michigan)
While the number of women who choose to pursue computer science and engineering careers is growing, men continue to largely outnumber them. In this paper, we describe a data mining approach that relies on a large collection of scientific articles to identify differences in gender interests in this field. Our hope is that through a better understanding of the differences between male and female preferences, we can enable more effective outreach and retention, and consequently contribute to the growth of the number of women who choose to pursue careers in this field.
I Spy: An Interactive Game-Based Approach to Multimodal Robot Learning
Parde, Natalie Paige (University of North Texas) | Papakostas, Michalis (University of Texas Arlington and NCSR Demokritos) | Tsiakas, Konstantinos (University of Texas Arlington and NCSR Demokritos) | Dagioglou, Maria (NCSR Demokritos) | Karkaletsis, Vangelis (NCSR Demokritos) | Nielsen, Rodney D (University of North Texas)
Teaching robots about objects in their environment requires a multimodal correlation of images and linguistic descriptions to build complete feature and object models. ย These models can be created manually by collecting images and related keywords and presenting the pairings to robots, but doing so is tedious and unnatural. ย This work abstracts the problem of training robots to learn about the world around them by introducing I Spy , an interactive dialogue- and vision-based game in which players place objects in front of a humanoid robot and challenge it to guess which object they have in mind. ย The robot gradually learns about the objects and the features which describe them through repeated games, by updating its knowledge with newly captured training images. ย This paper details I Spy's learning and gaming processes, describes the approaches taken to extract information from multiple modalities both before and during gameplay, and finally discusses the results of a study designed to evaluate the game's model accuracy over time, its overall performance, and its appeal to human players.
Modeling Spatial-Temporal Dynamics of Human Movements for Predicting Future Trajectories
Wang, Zhan (KTH Royal Institute of Technology) | Jensfelt, Patric (KTH Royal Institute of Technology) | Folkesson, John (KTH Royal Institute of Technology)
This paper presents a novel approach to modeling the dynamics of human movements with a grid-based representation.For each grid cell, we formulate the local dynamics using a variant of the left-to-right HMM, and thus explicitly model the exiting direction from the current cell. The dependency of this process on the entry direction is captured by employing the Input-Output HMM (IOHMM). On a higher level, we introduce the place where the whole trajectory originated into the IOHMM framework forming a hierarchical input structure. Therefore, we manage to capture both local spatial-temporal correlations and the long-term dependency on faraway initiating events, thus enabling the developed model to incorporate more information and to generate more informative predictions of future trajectories.The experimental results in an office corridor environment verify the capabilities of our method.
Interactive Multi-Consumer Power Cooperatives with Learning and Axiomatic Cost and Risk Disaggregation
Ehsanfar, Abbas (Stevens Institute of Technology) | Heydari, Babak (Stevens Institute of Technology)
This paper introduces a novel autonomous interactive learning cooperative (ILCP) who receives expected value and variance of load from consumers and participates in the electricity market on their behalf. Using an axiomatic approach, the share of each consumer's payment as well as its weight in calculating the modification of total day-ahead load are formulated. This scheme applies double-seasonal smoothing exponential, a recent load forecasting technique, and a classifier for real-time to day-ahead price direction forecasting (Gaussian Naรฏve Bayes). In addition to this, the ILCP employs interactive cooperative algorithms for both trading cooperative and consumer side. The ILCP scheme is investigated and its performance is compared to those of non-cooperative real-time pricing (RTP), LCP (non-interactive learning cooperative) and CP (non-interactive non-learning cooperative). The developed system was implemented using PJM(world's largest ย wholesale electricity market) real-time and day-ahead data for 2013 and half of 2014; real load profiles were selected from a set of 579 residential and commercial consumers, and weather data were applied to forecasting electricity price direction. We demonstrate the advantages of ILCP to lower the average electricity cost and to reduce unit price variations.
Termination Approximation: Continuous State Decomposition for Hierarchical Reinforcement Learning
Harris, Sean (University of New South Wales) | Hengst, Bernhard (University of New South Wales) | Pagnucco, Maurice (University of New South Wales)
This paper presents a divide-and-conquer decomposition for solving continuous state reinforcement learning problems. The contribution lies in a method for stitching together continuous state subtasks in a near-seamless manner along wide continuous boundaries. We introduce the concept of Termination Approximation where the set of subtask termination states are covered by goal sets to generate a set of subtask option policies. The approach employs hierarchical reinforcement learning methods and exploits any underlying repetition in continuous problems to allow reuse of the option policies both within a problem and across related problems. The approach is illustrated using a series of challenging racecar problems.
Teaching AI Ethics Using Science Fiction
Burton, Emanuelle (Center College) | Goldsmith, Judy (University of Kentucky) | Mattei, Nicholas (NICTA and University of New South Wales)
The cultural and political implications of modern AI research are not some far off concern, they are things that affect the world in the here and now. From advanced control systems with advanced visualizations and image processing techniques that drive the machines of the modern military to the slow creep of a mechanized workforce, ethical questions surround us. Part of dealing with these ethical questions is not just speculating on what could be but teaching our students how to engage with these ethical questions. We explore the use of science fiction as an appropriate tool to enable AI researchers to help engage students and the public on the current state and potential impacts of AI.
Unsupervised Domain Adaptation by Backpropagation
Ganin, Yaroslav, Lempitsky, Victor
Top-performing deep architectures are trained on massive amounts of labeled data. In the absence of labeled data for a certain task, domain adaptation often provides an attractive option given that labeled data of similar nature but from a different domain (e.g. synthetic images) are available. Here, we propose a new approach to domain adaptation in deep architectures that can be trained on large amount of labeled data from the source domain and large amount of unlabeled data from the target domain (no labeled target-domain data is necessary). As the training progresses, the approach promotes the emergence of "deep" features that are (i) discriminative for the main learning task on the source domain and (ii) invariant with respect to the shift between the domains. We show that this adaptation behaviour can be achieved in almost any feed-forward model by augmenting it with few standard layers and a simple new gradient reversal layer. The resulting augmented architecture can be trained using standard backpropagation. Overall, the approach can be implemented with little effort using any of the deep-learning packages. The method performs very well in a series of image classification experiments, achieving adaptation effect in the presence of big domain shifts and outperforming previous state-of-the-art on Office datasets.
Second-order Quantile Methods for Experts and Combinatorial Games
Koolen, Wouter M., van Erven, Tim
We aim to design strategies for sequential decision making that adjust to the difficulty of the learning problem. We study this question both in the setting of prediction with expert advice, and for more general combinatorial decision tasks. We are not satisfied with just guaranteeing minimax regret rates, but we want our algorithms to perform significantly better on easy data. Two popular ways to formalize such adaptivity are second-order regret bounds and quantile bounds. The underlying notions of 'easy data', which may be paraphrased as "the learning problem has small variance" and "multiple decisions are useful", are synergetic. But even though there are sophisticated algorithms that exploit one of the two, no existing algorithm is able to adapt to both. In this paper we outline a new method for obtaining such adaptive algorithms, based on a potential function that aggregates a range of learning rates (which are essential tuning parameters). By choosing the right prior we construct efficient algorithms and show that they reap both benefits by proving the first bounds that are both second-order and incorporate quantiles.
Plagiarism Detection in Polyphonic Music using Monaural Signal Separation
De, Soham, Roy, Indradyumna, Prabhakar, Tarunima, Suneja, Kriti, Chaudhuri, Sourish, Singh, Rita, Raj, Bhiksha
Given the large number of new musical tracks released each year, automated approaches to plagiarism detection are essential to help us track potential violations of copyright. Most current approaches to plagiarism detection are based on musical similarity measures, which typically ignore the issue of polyphony in music. We present a novel feature space for audio derived from compositional modelling techniques, commonly used in signal separation, that provides a mechanism to account for polyphony without incurring an inordinate amount of computational overhead. We employ this feature representation in conjunction with traditional audio feature representations in a classification framework which uses an ensemble of distance features to characterize pairs of songs as being plagiarized or not. Our experiments on a database of about 3000 musical track pairs show that the new feature space characterization produces significant improvements over standard baselines.