Technology
Policy Evaluation with Variance Related Risk Criteria in Markov Decision Processes
Tamar, Aviv, Di Castro, Dotan, Mannor, Shie
In this paper we extend temporal difference policy evaluation algorithms to performance criteria that include the variance of the cumulative reward. Such criteria are useful for risk management, and are important in domains such as finance and process control. We propose both TD(0) and LSTD(lambda) variants with linear function approximation, prove their convergence, and demonstrate their utility in a 4-dimensional continuous state space problem.
TRUSTS: Scheduling Randomized Patrols for Fare Inspection in Transit Systems Using Game Theory
Yin, Zhengyu (University of Southern California) | Jiang, Albert Xin (University of Southern California) | Tambe, Milind (University of Southern California) | Kiekintveld, Christopher (University of Texas at El Paso) | Leyton-Brown, Kevin (University of British Columbia) | Sandholm, Tuomas (Carnegie Mellon University) | Sullivan, John P. (Los Angeles County Sheriff's Department)
In proof-of-payment transit systems, passengers are legally required to purchase tickets before entering but are not physically forced to do so. Instead, patrol units move about the transit system, inspecting the tickets of passengers, who face fines if caught fare evading. TRUSTS models the problem of computing patrol strategies as a leader-follower Stackelberg game where the objective is to deter fare evasion and hence maximize revenue. We present an efficient algorithm for computing such patrol strategies and present experimental results using real-world ridership data from the Los Angeles Metro Rail system.
Reports of the AAAI 2012 Conference Workshops
Agrawal, Vikas (Infosys Limited) | Baier, Jorge (Pontificia Universidad Católica de Chile) | Bekris, Kostas (Rutgers University) | Chen, Yiling (Harvard University) | Garcez, Artur S. d'Avila (City University London,) | Hitzler, Pascal (Wright State University) | Haslum, Patrik (Australian National University) | Jannach, Dietmar (TU Dortmund) | Law, Edith (Carnegie Mellon University) | Lecue, Freddy (IBM Research) | Lamb, Luis C. (Federal University of Rio Grande do Sul) | Matuszek, Cynthia (University of Washington) | Palacios, Hector (Universidad Carlos III de Madrid) | Srivastava, Biplav (IBM Research) | Shastri, Lokendra (Infosys Limited) | Sturtevant, Nathan (University of Denver) | Stern, Roni (Ben Gurion University of the Negev) | Tellex, Stefanie (Massachusetts Institute of Technology) | Vassos, Stavros (National and Kapodistrian University of Athens)
The Multi-Agent Programming Contest
Behrens, Tristan (Clausthal University of Technology) | Dastani, Mehdi (Utrecht University) | Dix, Jürgen (Clausthal University of Technology) | Hübner, Jomi (University of Santa Catarina) | Köster, Michael (Clausthal University of Technology) | Novák, Peter (Delft University of Technology) | Schlesinger, Federico (Clausthal University of Technology)
The international Multi-Agent Programming Contest (MAPC), is a community-serving effort to facilitate advances in programming multiagent systems (MAS) by (1) developing benchmark problems, (2) enabling head-to-head comparison of MAS's and (3) supporting educational efforts in the design and implementation of MAS's.
Decision Making in Complex Multiagent Contexts: A Tale of Two Frameworks
Doshi, Prashant J. (University of Georgia)
Decision making is a key feature of autonomous systems. The physical context often includes other interacting autonomous systems, typically called agents. In this article, I focus on decision making in a multiagent context with partial information about the problem. I put the two frameworks, decentralized partially observable Markov decision process (Dec-POMDP) and the interactive partially observable Markov decision process (I-POMDP), in context and review the foundational algorithms for these frameworks, while briefly discussing the advances in their specializations.
The Answer Set Programming Competition
Calimeri, Francesco (Universita') | Ianni, Giovambattista (della Calabria) | Krennwallner, Thomas (Universita') | Ricca, Francesco (della Calabria)
The Answer Set Programming (ASP) Competition is a biannual event for evaluating declarative knowledge representation systems on hard and demanding AI problems. The competition consists of two main tracks: the ASP system track and the model and solve track. The traditional system track compares dedicated answer set solvers on ASP benchmarks, while the model and solve track invites any researcher and developer of declarative knowledge representation systems to participate in an open challenge for solving sophisticated AI problems with their tools of choice. This article provides an overview of the ASP competition series, reviews its origins and history, giving insights on organizing and running such an elaborate event, and briefly discusses about the lessons learned so far.
PROTECT -- A Deployed Game Theoretic System for Strategic Security Allocation for the United States Coast Guard
An, Bo (University of Southern California) | Shieh, Eric (University of Southern California) | Tambe, Milind (University of Southern California) | Yang, Rong (University of Southern California) | Baldwin, Craig (United States Coast Guard) | DiRenzo, Joseph (United States Coast Guard) | Maule, Ben (United States Coast Guard) | Meyer, Garrett (United States Coast Guard)
While three deployed applications of game theory for security have recently been reported, we as a community of agents and AI researchers remain in the early stages of these deployments; there is a continuing need to understand the core principles for innovative security applications of game theory. PROTECT is premised on an attacker-defender Stackelberg game model and offers five key innovations. First, this system is a departure from the assumption of perfect adversary rationality noted in previous work, relying instead on a quantal response (QR) model of the adversary's behavior --- to the best of our knowledge, this is the first real-world deployment of the QR model. Fourth, our experimental results illustrate that PROTECT's QR model more robustly handles real-world uncertainties than a perfect rationality model.
Machine Learning for Personalized Medicine: Predicting Primary Myocardial Infarction from Electronic Health Records
Weiss, Jeremy C. (University of Wisconsin-Madison) | Natarajan, Sriraam (Wake Forest University) | Peissig, Peggy L. (Marshfield Clinic Research Foundation) | McCarty, Catherine A. (Essentia Institute of Rural Health) | Page, David (University of Wisconsin-Madison)
Electronic health records (EHRs) are an emerging relational domain with large potential to improve clinical outcomes. We apply two statistical relational learning (SRL) algorithms to the task of predicting primary myocardial infarction. We show that one SRL algorithm, relational functional gradient boosting, outperforms propositional learners particularly in the medically-relevant high recall region. We observe that both SRL algorithms predict outcomes better than their propositional analogs and suggest how our methods can augment current epidemiological practices.
Towards Adapting Cars to their Drivers
Rosenfeld, Avi (Jerusalem College of Technology) | Bareket, Zevi (University of Michigan) | Goldman, Claudia V. (General Motors) | Kraus, Sarit (Bar-Ilan University) | LeBlanc, David J. (University of Michigan) | Tsimhoni, Omer (General Motors)
Such interactive activity leads us to consider intelligent and advanced ways of interaction leading to cars that can adapt to their drivers.In this paper, we focus on the Adaptive Cruise Control (ACC) technology that allows a vehicle to automatically adjust its speed to maintain a preset distance from the vehicle in front of it based on the driver's preferences. We introduce a method to combine machine learning algorithms with demographic information and expert advice into existing automated assistive systems. This method can reduce the interactions between drivers and automated systems by adjusting parameters relevant to the operation of these systems based on their specific drivers and context of drive. While generic packages such as Weka were successful in learning drivers' behavior, we found that improved learning models could be developed by adding information on drivers' demographics and a previously developed model about different driver types.