Technology
Uninformed-to-Informed Exploration in Unstructured Real-World Environments
Axelrod, Allan (Oklahoma State University) | Chowdhary, Girish (Oklahoma State University)
Conventionally, the process of learning the model (exploration) is initialized as either an uninformed or informed policy, where the latter leverages observations to guide future exploration. Informed exploration is ideal as it may allow a model to be learned in fewer samples. However, informed exploration cannot be implemented from the onset when a-priori knowledge on the sensing domain statistics are not available; such policies would only sample the first set of locations, repeatedly. Hence, we present a theoretically-derived bound for transitioning from uninformed exploration to informed exploration for unstructured real-world environments which may be partially-observable and time-varying. This bound is used in tandem with a sparsified Bayesian nonparametric Poisson Exposure Process, which is used to learn to predict the value of information in partiallyobservable and time-varying domains. The result is an uninformed-to-informed exploration policy which outperforms baseline algorithms in real-world data-sets.
Bayesian Clustering of Player Styles for Multiplayer Games
Normoyle, Aline (University of Pennsylvania) | Jensen, Shane T. (The Wharton School, University of Pennsylvania)
Clustering is an essential game analysis tool for understanding There are many clustering procedures that could be used player strengths and preferences. For example, clustering to group players based upon their play styles, with k-means techniques have been used to identify player preferences clustering being the most common method. Our use of for using vehicles over direct combat (Drachen et al. 2012), a model-based semi-parametric Bayesian clustering procedure for taking time to solve puzzles over running through content has two important advantages. First, the number of (Drachen, Canossa, and Yannakakis 2009), for understanding clusters (unique player styles) does not have to be prespecified.
Maximizing Flow as a Metacontrol in Angband
Mariusdottir, Thorey Maria (University of Alberta) | Bulitko, Vadim (University of Alberta) | Brown, Matthew (University of Alberta)
Flow is a psychological state that is reported to improve peopleโs performance. Flow can emerge when the personโs skills and the challenges of their activity match. This paper applies this concept to artificial intelligence agents. We equip a decision-making agent with a metacontrol policy that guides the agent to activities where the agentโs skills match the activity difficulty. Consequently, we expect the agentโs performance to improve. We implement and evaluate this approach in the role-playing game of Angband.
Autonomous Electricity Trading Using Time-Of-Use Tariffs in a Competitive Market
Urieli, Daniel (The University of Texas at Austin) | Stone, Peter (The University of Texas at Austin)
This research studies the impact of Time-Of-Use (TOU) tariffs in a competitive electricity market place. Specifically, it focuses on the question of how should an autonomous broker agent optimize TOU tariffs in a competitive retail market, and what is the impact of such tariffs on the economy. We formalize the problem of TOU tariff optimization and propose an algorithm for approximating its solution. We extensively experiment with our algorithm in a large-scale, detailed electricity retail markets simulation of the Power Trading Agent Competition (Power TAC) and: 1) find that our algorithm results in 15\% peak-demand reduction, 2) find that its peak-flattening results in greater profits and/or profit-share for the broker and allows it to win in head-to-head competition against the 1st and 2nd place brokers from the Power TAC 2014 finals, and 3) analyze several economic implications of using TOU tariffs in competitive retail markets.
A Benchmark for StarCraft Intelligent Agents
Uriarte, Alberto (Drexel University) | Ontaรฑรณn, Santiago (Drexel University)
The problem of comparing the performance of different Real-Time Strategy (RTS) Intelligent Agents (IA) is non-trivial. And often different research groups employ different testing methodologies designed to test specific aspects of the agents. However, the lack of a standard process to evaluate and compare different methods in the same context makes progress assessment difficult. In order to address this problem, this paper presents a set of benchmark scenarios and metrics aimed at evaluating the performance of different techniques or agents for the RTS game StarCraft. We used these scenarios to compare the performance of a collection of bots participating in recent StarCraft AI (Artificial Intelligence) competitions to illustrate the usefulness of our proposed benchmarks.
OntoAgents Gauge Their Confidence In Language Understanding
McShane, Marjorie (Rensselaer Polytechnic Institute) | Nirenburg, Sergei (Rensselaer Polytechnic Institute)
This paper details how OntoAgents, language-endowed intelligent agents developed in the OntoAgent framework, assess their confidence in understanding language inputs. It presents scoring heuristics for the following subtasks of natural language understanding: lexical disambiguation and the establishment of semantic dependencies; reference resolution; nominal compounding; the treatment of fragments; and the interpretation of indirect speech acts. The scoring of confidence in individual linguistic subtasks is a prerequisite for computing the overall confidence in the understanding of an utterance. This, in turn, is a prerequisite for the agentโs deciding how to act upon that level of understanding.
Sampling Hyrule: Multi-Technique Probabilistic Level Generation for Action Role Playing Games
Summerville, Adam James (University of California, Santa Cruz) | Mateas, Michael (University of California, Santa Cruz)
Procedural Content Generation (PCG) using machine learning is a fast growing area of research. Action Role Playing Game (ARPG) levels represent an interesting challenge for PCG due to their multi-tiered structure and nonlinearity. Previous work has used Bayes Nets (BN) to learn properties of the topological structure of levels from The Legend of Zelda. In this paper we describe a method for sampling these learned distributions to generate valid, playable level topologies. We carry this deeper and learn a sampleable representation of the individual rooms using Principal Component Analysis. We combine the two techniques and present a multi-scale machine learned technique for procedurally generating ARPG levels from a corpus of levels from The Legend of Zelda.
A Unified Framework for Human-Robot Knowledge Transfer
Shukla, Nishant (University of California, Los Angeles) | Xiong, Caiming (University of California, Los Angeles) | Zhu, Song-Chun (University of California, Los Angeles)
Transferring knowledge is a vital skill between humans for efficiently learning a new concept. In a perfect system, a human demonstrator can teach a robot a new task by using natural language and physical gestures. The robot would gradually accumulate and refine its spatial, temporal, and causal understanding of the world. The knowledge can then be transferred back to another human, or further to another robot. The implications of effective human to robot knowledge transfer include the compelling opportunity of a robot acting as the teacher, guiding humans in new tasks. The technical difficulty in achieving a robot implementation Figure 1: The robot autonomously performs a cloth folding of this caliber involves both an expressive knowledge task after learning from a human demonstration.
Can Accomplices to Fraud Will Themselves to Innocence, and Thereby Dodge Counter-Fraud Machines?
Bringsjord, Selmer (Rensselaer Polytechnic Institute (RPI) | Bringsjord, Alexander (Deep Detection LLC)
This brief paper explores the consequences of agnosticism with respect to whether a given human agent B is guilty of fraud. We find that if a human A is agnostic with respect to whether a human fraudster B is guilty of fraud, A, on the only formal definition of fraud that we are aware of, is her/himself provably not guilty of fraud. This means that a counter-fraud machine D based on an implemented version of this definition will classify A as innocent. Hence, if A by simply an act of will can bring it about that A is agnostic, A will evade D