Country
Clustering via Dirichlet Process Mixture Models for Portable Skill Discovery
Niekum, Scott (University of Massachusetts Amherst) | Barto, Andrew G. (University of Massachusetts Amherst)
Skill discovery algorithms in reinforcement learning typically identify single states or regions in state space that correspond to potential task-specific subgoals. However, such methods do not directly address the question of how many distinct skills are appropriate for solving the tasks that the agent faces. This can be highly inefficient when many identified subgoals correspond to the same underlying skill, but are all used in- dividually as skill goals. Furthermore, skills created in this manner are often only transferable to tasks that share iden- tical state spaces, since corresponding subgoals across tasks are not merged into a single skill goal. We show that these problems can be overcome by clustering subgoal data defined in an agent-space and using the resulting clusters as templates for skill termination conditions. Clustering via a Dirichlet process mixture model is used to discover a minimal, suffi- cient collection of portable skills.
InfoMax Control for Acoustic Exploration of Objects by a Mobile Robot
Rebguns, Antons ( Department of Computer Sceince School of Information: Science, Technology, and Arts University of Arizona ) | Ford, Daniel ( Department of Electrical and Computer Engineering University of Arizona ) | Fasel, Ian R ( School of Information: Science, Technology, and Arts University of Arizona )
Recently, information gain has been proposed as a candidate intrinsic motivation for lifelong learning agents that may not always have a specific task. In the InfoMax control framework, reinforcement learning is used to find a control policy for a POMDP in which movement and sensing actions are selected to reduce Shannon entropy as quickly as possible. In this study, we implement InfoMax control on a robot which can move between objects and perform sound-producing manipulations on them. We formulate a novel latent variable mixture model for acoustic similarities and learn InfoMax polices that allow the robot to rapidly reduce uncertainty about the categories of the objects in a room. We find that InfoMax with our improved acoustic model leads to policies which lead to high classification accuracy. Interestingly, we also find that with an insufficient model, the InfoMax policy eventually learns to "bury its head in the sand" to avoid getting additional evidence that might increase uncertainty. We discuss the implications of this finding for InfoMax as a principle of intrinsic motivation in lifelong learning agents.
Improving Consensus Accuracy via Z-Score and Weighted Voting
Jung, Hyun Joon (University of Texas at Austin) | Lease, Matthew (University of Texas at Austin)
Using supervised and unsupervised features individually or together, we (a) detect and filter out noisy workers via Z-score, and (b) weight worker votes for consensus labeling. We evaluate on noisy labels from Amazon Mechanical Turk in which workers judge Web search relevance of query/document pairs. In comparison to a majority vote baseline, results show a 6% error reduction (48.83% to 51.91%) for graded accuracy and 5% error reduction (64.88% to 68.33%) for binary accuracy.
Using Gaussian Process Regression for Efficient Motion Planning in Environments with Deformable Objects
Frank, Barbara (Albert-Ludwigs-University, Freiburg) | Stachniss, Cyrill (Albert-Ludwigs-University, Freiburg) | Abdo, Nichola (Albert-Ludwigs-University, Freiburg) | Burgard, Wolfram (Albert-Ludwigs-University, Freiburg)
The ability to plan their own motions and to reliably execute them is an important precondition for autonomous robots. In this paper, we consider the problem of planning the motion of a mobile manipulation robot in the presence of deformable objects in the environment. Our approach combines probabilistic roadmap planning with a deformation simulation system. Since the physical deformation simulation is computationally demanding, we use an efficient variant of Gaussian process regression to estimate the deformation cost for individual objects based on training examples. We generate the training data by employing a simulation system in a preprocessing step. Consequently, no simulations are needed during runtime. We implemented and tested our approach on a mobile manipulation robot. Our experiments show that the robot is able to accurately predict and thus consider the deformation cost its manipulator introduces to the environment during motion planning. Simultaneously, the computation time is substantially reduced compared to a system that performs physical simulations online.
InSitu: An Approach for Dynamic Context Labeling Based on Product Usage and Sound Analysis
Lyardet, Fernando (Technical University of Darmstadt) | Hadjakos, Aristotelis (Technical University of Darmstadt) | Szeto, Diego Wong (Technical University of Darmstadt)
Smart environments offer a vision of unobtrusive interaction with our surroundings, interpreting and anticipating our needs. One key aspect for making environments smart is the ability to recognize the current context. However, like any human space, smart environments are subject to changes and mutations of their purposes and their composition as people shape their living places according to their needs. In this paper we present an approach for recognizing context situations in smart environments that addresses this challenge. We propose a formalism for describing and sharing context states (or situations) and an architecture for gradually introducing contextual knowledge to an environment, where the current context is determined on sensing people's usage of devices and sound analysis.
Continual HTN Robot Task Planning in Open-Ended Domains: A Case Study
Off, Dominik (University of Hamburg) | Zhang, Jianwei (University of Hamburg)
The fact that many AI planning approaches are still based on too simplifying assumptions makes it often hard to apply these approaches to real-world robotics. In particular, it is in many cases difficult to generate a complete plan in advance, because not all information is available at the beginning of the planning process. We briefly present the continual planning system ACogPlan and a preliminary test case that demonstrates how the planning system can enable mobile robots to continually plan and execute activities in an open-ended domain.
Efficiently Eliciting Preferences from a Group of Users
Hines, Greg (University of Waterloo) | Larson, Kate (University of Waterloo)
Learning about users' preferences allows agents to make intelligent decisions on behalf of users. When we are eliciting preferences from a group of users, we can use the preferences of the users we have already processed to increase the efficiency of the elicitation process for the remaining users. However, current methods either require strong prior knowledge about the users' preferences or can be overly cautious and inefficient. Our method, based on standard techniques from non-parametric statistics, allows the controller to choose a balance between prior knowledge and efficiency. This balance is investigated through experimental results.
Leading Multiple Ad Hoc Teammates in Joint Action Settings
Agmon, Noa (The University of Texas at Austin) | Stone, Peter (The University of Texas at Austin)
The growing use of autonomous agents in practice may require agents to cooperate as a team in situations where they have limited prior knowledge about one another, cannot communicate directly, or do not share the same world models. These situations raise the need to design ad hoc team members, i.e., agents that will be able to cooperate without coordination in order to reach an optimal team behavior. This paper considers problem of leading N-agent teams by a single agent toward their optimal joint utility, where the agents compute their next actions based only on their most recent observations of their teammates' actions. We show that compared to previous results in two-agent teams, in larger teams the agent might not be able to lead the team to the action with maximal joint utility. In these cases, the agent's optimal strategy leads the team to the best possible reachable cycle of joint actions. We describe a graphical model of the problem and a polynomial time algorithm for solving it. We then consider the problem of leading teams where the agents' base their actions on a longer history of past observations, showing that the an upper bound computation time exponential in the memory size is very likely to be tight.
A Network View of Human Ingestion and Health: Instrumental Artificial Intelligence
Edgell, Robert Anthony (American University) | Vogl, Roland (Stanford University)
Humans are confronted with an increasingly complex array of ingestion substances and dietary choices that influence health and well being. However, even with strong medical evidence that clearly links ingestion strategies and heath consequences, the general public struggles to make health-optimizing ingestion decisions. Based on our literature review, we delineate a typology of barriers to formulating health-optimizing ingestion strategies. We propose that the introduction of artificial intelligence (AI) as “decision management” (AI-DM) technology into the ingestion decision-making network would increase the likelihood of more predictable and optimized health outcomes. Also, we delineate the key informational constituencies needed to enable a comprehensive and effective AI-DM system. While no author has yet proposed AI in the particular context discussed in this paper, the theoretical and empirical literature suggests that this might be possible. We conclude by discussing areas for additional research.
Improvement of Multi-AUV Cooperation through Teammate Verification
Novitzky, Michael (The Georgia Institute of Technology)
Current methods for multi-AUV cooperation suffer in low communication environments. State of the art methods employ auctioneering or planning to determine a single AUV'task. These systems require communication to update models of teammates and tasks for efficient task selection. Most strategies assume a teammate is inoperable if a communication timeout is reached which reduces overall team efficiency. Including teammate prediction has been shown to mitigate efficiency degeneration due to low communication. However, there is no verification of a predicted teammate's task other than through eventual communication. A possible verification tool is behavior recognition. Current behavior recognition utilizes either overhead sensors or post mission analysis to track robot trajectories in order to infer their internal state. A system in which an AUV is capable of sensing a teammate, for example through a forward-looking sonar, and deducing it's behavior along with contextual information, such as location, will enable an AUV to determine that teammate's current task in the overall mission. This will allow for an accurate update of that teammate's model allowing the AUV to more efficiently determine its own next task rather than relying only on communication. This position paper posits that multi-AUV cooperation efficiency will improve in low communication environments with the combination of robust teammate prediction along with verification using behavior recognition.