Country
Integrating Opponent Models with Monte-Carlo Tree Search in Poker
Ponsen, Marc (Maastricht University) | Gerritsen, Geert (Maastricht University) | Chaslot, Guillaume (Maastricht University)
In this paper we apply a Monte-Carlo Tree Search implementation that is boosted with domain knowledge to the game of poker. More specifically, we integrate an opponent model in the Monte-Carlo Tree Search algorithm to produce a strong poker playing program. Opponent models allow the search algorithm to focus on relevant parts of the game-tree. We use an opponent modelling approach that starts from a (learned) prior, i.e., general expectations about opponent behavior, and then learns a relational regression tree-function that adapts these priors to specific opponents. Our modelling approach can generate detailed game features or relations on-the-fly. Additionally, using a prior we can already make reasonable predictions even when limited experience is available for a particular player. We show that Monte-Carlo Tree Search with integrated opponent models performs well against state-of-the-art poker programs.
Declarative Probabilistic Programming for Undirected Graphical Models: Open Up to Scale Up
Riedel, Sebastian Robert (University of Massachusetts)
We argue that probabilistic programming with undirected models, in order to scale up, needs to open up. That is, instead of focusing on minimal sets of generic building blocks such as universal quantification or logical connectives, languages should grow to include specific building blocks for as many uses cases as necessary. This can not only lead to more concise models, but also to more efficient inference if we use methods that can exploit building-block specific sub-routines. As embodiment of this paradigm we present , a platform for implementing probabilistic programming languages that grow.
Multiagent Meta-Level Control for Predicting Meteorological Phenomena
Cheng, Shanjun (The University of North Carolina at Charlotte) | Raja, Anita (The University of North Carolina at Charlotte) | Lesser, Victor (University of Massachusetts Amherst)
It is crucial for social systems to adapt to the dynamics of open environments. This adaptation process becomes especially challenging in the context of multiagent systems. In this paper, we argue that multiagent meta-level control is an effective way to determine when this adaptation process should be done and how much effort should be invested in adaptation as opposed to continuing with the current action plan. We develop a reinforcement learning based mechanism for multiagent meta-level control that facilitates the metalevel control component of each agent to learn policies in a decentralized fashion that (a) it can efficiently support agent interactions with other agents and (b) reorganize the underlying network when needed. We evaluate this mechanism in the context of a multiagent tornado tracking application called NetRads. Empirical results show that adaptive multiagent meta-level control significantly improves the performance of the tornado tracking network for a variety of weather scenarios.
Machine Reading: A "Killer App" for Statistical Relational AI
Poon, Hoifung (University of Washington) | Domingos, Pedro (University of Washington)
Machine reading aims to automatically extract knowledge from text. It is a long-standing goal of AI and holds the promise of revolutionizing Web search and other fields. In this paper, we analyze the core challenges of machine reading and show that statistical relational AI is particularly well suited to address these challenges. We then propose a unifying approach to machine reading in which statistical relational AI plays a central role. Finally, we demonstrate the promise of this approach by presenting OntoUSP, an end-to-end machine reading system that builds on recent advances in statistical relational AI and greatly outperforms state-of-the-art systems in a task of extracting knowledge from biomedical abstracts and answering questions.
Speculations on Leveraging Graphical Models for Architectural Integration of Visual Representation and Reasoning
Rosenbloom, Paul (University of Southern California)
The starting point is an ongoing effort to structure underlying intelligent behavior, whether intended reconstruct cognitive architectures from the ground up via as models of human intelligence and/or implementations of graphical models (Koller and Friedman 2009), with the artificial intelligence (Langley, Laird and Rogers 2009). A aim of understanding existing architectures better, basic cognitive architecture may comprise memories, exploring the overall space of architectures, and decision algorithms, learning mechanisms, and some developing new and improved architectures (Rosenbloom means of interacting with external environments.
Metacognition for Detecting and Resolving Conflicts in Operational Policies
Josyula, Darsana (Bowie State University) | Donahue, Bette (Bowie State University) | McCaslin, Matthew (Bowie State University) | Snowden, Michelle (Franklin and Marshall College) | Anderson, Michael (University of Maryland Baltimore County) | Oates, Timothy (University of Maryland Baltimore County) | Schmill, Matthew (University of Maryland, College Park) | Perlis, Donald
Informational conflicts in operational policies cause agents to run into situations where responding based on the rules in one policy violates the same or another policy. Static checking of these conflicts is infeasible and impractical in a dynamic environment. This paper discusses a practical approach to handling policy conflicts in real-time domains within the context of a hierarchical military command and control simulated system that consists of a central command, squad leaders and squad members. All the entities in the domain function according to preset communication and action protocols in order to perform successful missions. Each entity in the domain is equipped with an instance of a metacognitive component to provide on-board/on-time analysis of actions and recommendations during the operation of the system. The metacognitive component is the Metacognitive Loop (MCL) which is a general purpose anomaly processor designed to function as a cross-domain plugin system. It continuously monitors expectations and notices when they are violated, assesses the cause of the violation and guides the host system to an appropriate response. MCL makes use of three ontologiesโindications, failures and responsesโto perform the notice, assess and guide phases when a conflict occurs. Conflicts in the set of rules (within a policy or between policies) manifest as expectation violations in the real world. These expectation violations trigger nodes in the indication ontology which, in turn, activate associated nodes in the failure ontology. The responding failure nodes then activate the appropriate nodes in the response ontology. Depending on which response node gets activated, the actual response may vary from ignoring the conflict to prioritizing, modifying or deleting one or more conflicting rules.
Motion Planning Algorithms for Autonomous Intersection Management
Au, Tsz-Chiu (The University of Texas at Austin) | Stone, Peter (The University of Texas at Austin)
The impressive results of the 2007 DARPA Urban Challenge showed that fully autonomous vehicles are technologically feasible with current intelligent vehicle hardware. It is natural to ask how current transportation infrastructure can be improved when most vehicles are driven autonomously in the future. Dresner and Stone proposed a new intersection control mechanism called Autonomous Intersection Management (AIM) and showed in simulation that intersection control can be made more efficient than the traditional control mechanisms such as traffic signals and stop signs. In this paper, we extend the study by examining the relationship between the precision of cars' motion controllers and the efficiency of the intersection controller. We propose a planning-based motion controller that can reduce the chance that autonomous vehicles stop before intersections, and show that this controller can increase the efficiency of the intersection control mechanism.
MCRNR: Fast Computing of Restricted Nash Responses by Means of Sampling
Ponsen, Marc (Maastricht University) | Lanctot, Marc (University of Alberta) | Jong, Steven de (Maastricht University)
This paper presents a sample-based algorithm for the computation of restricted Nash strategies in complex extensive form games. Recent work indicates that regret-minimization algorithms using selective sampling, such as Monte-Carlo Counterfactual Regret Minimization (MCCFR), converge faster to Nash-equilibrium (NE) strategies than their non-sampled counterparts which perform a full tree traversal. In this paper, we show that MCCFR is also able to establish NE strategies in the complex domain of Poker. Although such strategies are defensive (i.e. safe to play), they are oblivious to opponent mistakes. We can thus achieve better performance by using (an estimation of) opponent strategies. The Restricted Nash Response (RNR) algorithm was proposed to learn robust counter-strategies given such knowledge. It solves a modified game, wherein it is assumed that opponents play according to a fixed strategy with a certain probability, or to a regret-minimizing strategy otherwise. We improve the rate of convergence of the RNR algorithm using sampling. Our new algorithm, MCRNR, samples only relevant parts of the game tree. It is therefore able to converge faster to robust best-response strategies than RNR.We evaluate our algorithm on a variety of imperfect information games that are small enough to solve yet large enough to be strategically interesting, as well as a large game, Texas Holdโem Poker.
Parallel Best-First Search: The Role of Abstraction
Burns, Ethan (University of New Hampshire) | Lemons, Sofia (University of New Hampshire) | Ruml, Wheeler (University of New Hampshire) | Zhou, Rong (Palo Alto Research Center)
To harness modern multicore processors, it is imperative to develop parallel versions of fundamental algorithms. In this paper, we present a general approach to best-first heuristic search in a shared-memory setting. Each thread attempts to expand the most promising nodes. By using abstraction to partition the state space, we detect duplicate states while avoiding lock contention. We allow speculative expansions when necessary to keep threads busy. We identify and fix potential livelock conditions. In an empirical comparison on STRIPS planning, grid pathfinding, and sliding tile puzzle problems using an 8-core machine, we show that A* implemented in our framework yields faster search performance than previous parallel search proposals. We also demonstrate that our approach extends easily to other best-first searches, such as weighted A* and anytime heuristic search.
Using a Trust Model in Decision Making for Supply Chain Management
Haghpanah, Yasaman (University of Maryland, Baltimore County) | desJardins, Marie (University of Maryland, Baltimore County)
One of the critical factors for a successful cooperative relationship in a supply chain partnership is trust. Many real-world applications, such as Supply Chain Management (SCM), can be modeled using multi-agent systems. One shortcoming of current SCM models is that their trust models are ad hoc and do not have a strong theoretical basis. As a result, they are unable to model subtleties in agent behavior that can be used to build a more accurate trust model. We propose a trust model for SCM that is grounded in probabilistic game theory. In this model, trust can be gained through direct interactions and/or by asking for information from other trustworthy agents. We will use this model to simulate and study supply chain market behavior.