Agents
FAQ-Learning in Matrix Games: Demonstrating Convergence Near Nash Equilibria, and Bifurcation of Attractors in the Battle of Sexes
Kaisers, Michael (Maastricht University) | Tuyls, Karl (Maastricht University)
This article studies Frequency Adjusted Q-learning (FAQ-learning), a variation of Q-learning that simulates simultaneous value function updates. The main contributions are empirical and theoretical support for the convergence of FAQ-learning to attractors near Nash equilibria in two-agent two-action matrix games.The games can be divided into three types: Matching pennies, Prisoners' Dilemma and Battle of Sexes. This article shows that the Matching pennies and Prisoners' Dilemma yield one attractor of the learning dynamics, while the Battle of Sexes exhibits a supercritical pitchfork bifurcation at a critical temperature, where one attractor splits into two attractors and one repellent fixed point. Experiments illustrate that the distance between fixed points of the FAQ-learning dynamics and Nash equilibria tends to zero as the exploration parameter of FAQ-learning approaches zero.
Lifelong Credit Assignment with the Success-Story Algorithm
Schmidhuber, Juergen (The Swiss AI Lab IDSIA, University of Lugano, and SUPSI)
Consider an embedded agent with a self-modifying, Turing-equivalent policy that can change only through active self-modifications. How can we make sure that it learns to continually accelerate reward intake? Throughout its life the agent remains ready to undo any self-modification generated during any earlier point of its life, provided the reward per time since then has not increased, thus enforcing a lifelong success-story of self-modifications, each followed by long-term reward acceleration up to the present time. The stack-based method for enforcing this is called the success-story algorithm. It fully takes into account that early self-modifications set the stage for later ones (learning a learning algorithm), and automatically learns to extend self-evaluations until the collected reward statistics are reliable ... a very simple but general method waiting to be re-discovered! Time permitting, I will also briefly discuss more recent mathematically optimal universal maximizers of lifelong reward, in particular, the fully self-referential Goedel machine.
Agent Based Intelligent Decluttering Enhancements
Pfautz, Stacy Lovell (Aptima, Inc.) | Schurr, Nathan (Aptima, Inc.) | Ganberg, Gabriel (Aptima, Inc.) | Bauer, David (Aptima, Inc.) | Scerri, Paul (Carnegie Mellon University)
Model-driven visualization (MDV) is a novel framework that supports more effective, intelligent user interfaces to improve decision making in complex environments by coupling cognitive and perceptual theories of information processing with advanced artificial intelligence methods. It embeds empirical and theory driven approaches for identifying and prioritizing data based on the information requirements and needs of the human decision maker within intelligent agents. The agents automatically deliver and present information based on its likely value using visualizations that best convey that information to the user(s) of the system. Agents also reason about the context and constraints of the user, environment, and display to enable a higher degree of personalization within an interactive user interface (e.g., by drawing a userโs attention to interesting aspects of the data such as trends, anomalies, and patterns). We apply cognitive systems engineering processes to help identify the information available to individuals and/or teams, where it resides, where it is needed, and ultimately how to create the mappings required in connecting critical information to those who need it with innovative visualizations that most effectively support the end user. This paper describes the application of MDV to intelligently deliver timely, mission-critical information by adapting a Common Tactical Picture (CTP) display used for maritime situation awareness, threat assessment, and decision support.
Leadership Games and their Application in Super-Peer Networks
Walsh, Thomas John (University of Arizona) | Taheri, Javad (University of Arizona) | Wright, Jeremy Bryan (University of Arizona) | Cohen, Paul (University of Arizona)
This paper considers a setting where a single ``leadership agent'' intervenes in a multi-agent system through actions that (perhaps subtly) change the dynamics of the system. We describe a number of forms this intervention can take and compare these situations to settings in previous work. We identify two important effects of leadership: faster system convergence, and convergence to a better equilibrium. Empirically, we first explore these properties in leadership of algorithms engaged in classical 2-player games. We then apply this general framework to the leadership of a super-peer file-sharing network. In these experiments the network contains some agents that make locally greedy decisions that hamper the network as a whole. We show that a leader acting based on a more global criteria can push the system to a better equilibrium point as well as speeding up convergence. We also show how a mathematical approximation of such super-peer networks can be used to aid a leader in determining a minimum-cost intervention strategy.
Role-Based Ad Hoc Teamwork
Genter, Katie (University of Texas at Austin) | Agmon, Noa (University of Texas at Austin) | Stone, Peter (University of Texas at Austin)
An ad hoc team setting is one in which teammates must work together to obtain a common goal, but without any prior agreement regarding how to work together. In this paper we present a role-based approach for ad hoc teamwork, in which each teammate is inferred to be following a specialized role that accomplishes a specific task or exhibits a particular behavior. In such cases, the role an ad hoc agent should select depends both on its own capabilities and on the roles currently selected by the other team members. We formally define methods for evaluating the influence of the ad hoc agent's role selection on the team's utility, leading to an efficient calculation of the role that yields maximal team utility. In simple teamwork settings, we demonstrate that the optimal role assignment can be easily determined. However, in complex environments, where it is not trivial to determine the optimal role assignment, we examine empirically the best suited method for role assignment. Finally, we show that the methods we describe have a predictive nature. As such, once an appropriate assignment method is determined for a domain, it can be used successfully in new tasks that the team has not encountered before and for which only limited prior experience is available.
Dynamic User Task Scheduling for Mobile Robots
Coltin, Brian (Carnegie Mellon University) | Veloso, Manuela (Carnegie Mellon University) | Ventura, Rodrigo (Institute Superior Tecnico)
We present our efforts to deploy mobile robots in office environments, focusing in particular on the challenge of planning a schedule for a robot to accomplish user-requested actions. We concretely aim to make our CoBot mobile robots available to execute navigational tasks requested by users, such as telepresence, and picking up and delivering messages or objects at different locations. We contribute an efficient web-based approach in which users can request and schedule the execution of specific tasks. The scheduling problem is converted to a mixed integer programming problem. The robot executes the scheduled tasks using a synthetic speech and touch-screen interface to interact with users, while allowing users to follow the task execution online. Our robot uses a robust Kinect-based safe navigation algorithm, moves fully autonomously without the need to be chaperoned by anyone, and is robust to the presence of moving humans, as well as non-trivial obstacles, such as legged chairs and tables. Our robots have already performed 15km of autonomous service tasks.
Reciprocal Preference Model for Two Player Dilemma Games
Ahmed, Asrar (IIIT Hyderabad) | Karlapalem, Kamalakar (IIIT Hyderabad)
Results from behavioral economics show that individuals do not always maximize monetary payoffs. Within behavioral economics different models of social preference have been put forth to account for this deviation from standard assumptions of game theory and economics. Incorporating such models into agent decision making is increasingly relevant to design systems which interact with or on behalf of humans. Existing models, which correctly predict outcomes across a large set of games, are fairly complex. To this end, we present aspiration based social preference model and evaluate it by considering two player dilemma games. We show that the qualitative predictions of our model are consistent with results from behavioral economics.
Detecting and Identifying Coalitions
Kerr, Reid (University of Waterloo) | Cohen, Robin (University of Waterloo)
In many multiagent scenarios, groups of participants (known as coalitions) may attempt to cooperate, seeking to increase the benefits realized by the members. Depending on the scenario, such cooperation may be benign, or may be unwelcome or even forbidden (often called collusion). Coalitions can present a problem for many multiagent systems, potentially undermining the intended operation of systems. In this paper, we present a technique for detecting the presence of coalitions (malicious or otherwise), and identifying their members. Our technique employs clustering in benefit space, a high-dimensional feature space reflecting the benefit flowing between agents, in order to identify groups of agents who are similar in terms of the agents they are favoring. A statistical approach is then used to characterize candidate clusters, identifying as coalitions those groups that favor their own members to a much greater degree than the general population. We believe that our approach is applicable to a wide range of domains. Here, we demonstrate its effectiveness within a simulated marketplace making use of a trust and reputation system to cope with dishonest sellers. Many trust and reputation proposals readily acknowledge their ineffectiveness in the face of collusion, providing one example of the importance of the problem. While certain aspects of coalitions have received significant attention (e.g., formation, stability, etc.), relatively little research has focused on the problem of coalition identification. We believe our research represents an important step towards addressing the challenges posed by coalitions.
Mechanism Design for Aggregated Demand Prediction in the Smart Grid
Rose, Harry Thomas (University of Southampton) | Rogers, Alex (University of Southampton) | Gerding, Enrico H (University of Southampton)
This paper presents a novel scoring rule-based mechanism that encourages agents to produce costly estimates of future events and truthfully report them to a centre when the budget for payments to the agents is itself determined by their reports. This is applied to a model of aggregated demand prediction within a microgrid where, given estimates of future consumptions, an aggregator must optimally purchase electricity for a set of homes, each represented by self-interested, rational home agents. This in turn reduces the need for costly standby generation within the grid. The aggregator has prior information about the amount each home will consume, and determines the amount to pay each agent based on savings resulting from using the agents' reported information, over its own prior information. Agents use sensory information regarding their property and its occupants to generate these estimates, which they transmit to the aggregator using smart grid technology. The proposed mechanism is dominant strategy incentive compatible and empirical evaluation shows that it encourages agents to exert effort in producing precise estimates. We show that the mechanism is ex ante individually rational for the aggregator, and that it outperforms a simpler mechanism whereby savings are distributed evenly.
Representing Context Using the Context for Human and Automation Teams Model
Ganberg, Gabriel (Aptima, Inc.) | Ayers, Jeanine (Aptima, Inc.) | Schurr, Nathan (Aptima, Inc.) | Therrien, Michael (Aptima, Inc.) | Rousseau, Jeff (Aptima, Inc.)
The goal of representing context in a mixed initiative sys-tem is to model the information at a level of abstraction that is actionable for both the human and automated system. A potential solution to this problem is the Context for Human and Automation Teams (CHAT). This paper introduces the CHAT model and provides example implementations from several different applications such as task scheduling tech-niques, multi-agent systems, and human-robot interaction.