Technology
An Empirical Evaluation of Evaluation Metrics of Procedurally Generated Mario Levels
Mariño, Julian R. H. (Universidade Federal de Viçosa) | Reis, Willian M. P. (Universidade Federal de Viçosa) | Lelis, Levi H. S. (Universidade Federal de Viçosa)
There are several approaches in the literature for automatically generating Infinite Mario Bros levels. The evaluation of such approaches is often performed solely with computational metrics such as leniency and linearity. While these metrics are important for an initial exploratory evaluation of the content generated, it is not clear whether they are able to capture the player's perception of the content generated. In this paper we evaluate several of the commonly used computational metrics. Namely, we perform a systematic user study with procedural content generation systems and compare the insights gained from our user study with those gained from analyzing the computational metric values. The results of our experiment suggest that current computational metrics should not be used in lieu of user studies for evaluating content generated by computer programs.
Guardian: A Crowd-Powered Spoken Dialog System for Web APIs
Huang, Ting-Hao Kenneth (Carnegie Mellon University) | Lasecki, Walter S. (University of Michigan) | Bigham, Jeffrey P. (Carnegie Mellon University)
Natural language dialog is an important and intuitive way for people to access information and services. However, current dialog systems are limited in scope, brittle to the richness of natural language, and expensive to produce. This paper introduces Guardian, a crowd-powered framework that wraps existing Web APIs into immediately usable spoken dialog systems. Guardian takes as input the Web API and desired task, and the crowd determines the parameters necessary to complete it, how to ask for them, and interprets the responses from the API. The system is structured so that, over time, it can learn to take over for the crowd. This hybrid systems approach will help make dialog systems both more general and more robust going forward.
Exploring Player Trace Segmentation for Dynamic Play Style Prediction
Valls-Vargas, Josep (Drexel University) | Ontañón, Santiago (Drexel University) | Zhu, Jichen (Drexel University)
Existing work on player modeling often assumes that the play style of players is static. However, our recent work shows evidence that players regularly change their play style over time. In this paper we propose a novel player modeling framework to capture this change by using episodic information and sequential machine learning techniques. In particular, we experiment with different trace segmentation strategies for play style prediction. We evaluate this new framework on gameplay data gathered from a game-based interactive learning environment. Our results show that sequential machine learning techniques that incorporate predictions from previous segments outperform non-sequential techniques. Our results also show that too fine (minute-by-minute) or too coarse (whole trace) segmentation of traces decreases performance.
Linguistic Wisdom from the Crowd
Chang, Nancy (Google) | Lee-Goldman, Russell (Google) | Tseng, Michael (Google)
Crowdsourcing for linguistic data typically aims to replicate expert annotations using simplified tasks. But an alternative goal — one that is especially relevant for research in the domains of language meaning and use — is to tap into people's rich experience as everyday users of language. Research in these areas has the potential to tell us a great deal about how language works, but designing annotation frameworks for crowdsourcing of this kind poses special challenges. In this paper we define and exemplify two approaches to linguistic data collection corresponding to these differing goals (model-driven and user-driven) and discuss some hybrid cases in which they overlap. We also describe some design principles and resolution techniques helpful for eliciting linguistic wisdom from the crowd.
The ActiveCrowdToolkit: An Open-Source Tool for Benchmarking Active Learning Algorithms for Crowdsourcing Research
Venanzi, Matteo (University of Southampton) | Parson, Oliver (University of Southampton) | Rogers, Alex (University of Southampton) | Jennings, Nick (University of Southampton)
Figure Crowdsourcing systems are commonly faced with the challenge 1 (a) shows the interface which allows researchers to set of making online decisions by assigning tasks to workers up experiments which run multiple active learning strategies in order to maximise accuracy while also minimising over a single dataset. Using this dialog, the user can construct cost. To aid researchers to reproduce, benchmark and extend an active learning strategy by combining an aggregation state-of-the-art active learning methods for crowdsourcing model, a task selection method and a worker selection systems, we developed the open-source.NET ActiveCrowd-method. The user can also select the number of judgements Toolkit.
Learning Supervised Topic Models from Crowds
Rodrigues, Filipe (University of Coimbra) | Ribeiro, Bernardete (University of Coimbra) | Lourenço, Mariana (University of Coimbra) | Pereira, Francisco (Massachusetts Institute of Technology)
The growing need to analyze large collections of documents has led to great developments in topic modeling. Since documents are frequently associated with other related variables, such as labels or ratings, much interest has been placed on supervised topic models. However, the nature of most annotation tasks, prone to ambiguity and noise, often with high volumes of documents, deem learning under a single-annotator assumption unrealistic or unpractical for most real-world applications. In this paper, we propose a supervised topic model that accounts for the heterogeneity and biases among different annotators that are encountered in practice when learning from crowds. We develop an efficient stochastic variational inference algorithm that is able to scale to very large datasets, and we empirically demonstrate the advantages of the proposed model over state of the art approaches.
How Effective an Odd Message Can Be: Appropriate and Inappropriate Topics in Speech-Based Vehicle Interfaces
Sirkin, David (Stanford University) | Fischer, Kerstin (Southern Denmark University) | Jensen, Lars (Southern Denmark University) | Ju, Wendy (Stanford University and California College of the Arts)
Dialog between drivers and speech-based vehicle interfaces can be used as an instrument to find out what drivers might be concerned, confused or curious about in driving simulator studies. Eliciting on-going conversation with drivers about topics that go beyond navigation, control of entertainment systems, or other traditional driving related tasks is important to getting drivers to engage with the activity in an open-ended fashion. In a structured improvisational Wizard of Oz study that took place in a highly immersive driving simulator, we engaged participant drivers (N=6) in an autonomous driving course where the vehicle spoke to drivers using computer-generated natural language speech. Using microanalyses of the drivers’ responses to the car’s utter- ances, we identify a set of topics that are expected and treated as appropriate by the participants in our study, as well as a set of topics and conversational strategies that are treated as inappropriate. We also show that it is just these unexpected, inappropriate utterances that eventually increase users’ trust in the system, make them more at ease, and raise the system’s acceptability as a communication partner.
It’s Not Just What You Say, But How You Say It: Muiltimodal Sentiment Analysis Via Crowdsourcing
Elshenawy, Ahmad Khamis (University of Washington) | Carter, Steele (University of Washington) | Braga, Daniela (Voicebox Technologies)
This paper examines the effect of various modalities of expression on the reliability of crowdsourced sentiment polarity judgments. A novel corpus of YouTube video reviews was created, and sentiment judgments were obtained via Amazon Mechanical Turk. We created a system for isolating text, video, and audio modalities from YouTube videos to ensure that annotators could only see the particular modality or modalities being evaluated. Reliability of judgments was assessed using Fleiss Kappa inter-annotator agreement values. We found that the audio only modality produced the most reliable judgments for video fragments and that across modalities video fragments are less ambiguous than full videos.
Evaluating the Pairwise Event Salience Hypothesis in Indexter
Kives, Christopher (University of New Orleans) | Ware, Stephen G. (University of New Orleans) | Baker, Lewis J. (Vanderbilt University)
Indexter is a plan-based computational model of narrative discourse which leverages cognitive scientific theories of how events are stored in memory during online comprehension. These discourse models are valuable for static and interactive narrative generation systems because they allow the author to reason about the audience's understanding and attention as they experience a story. A pair of Indexter events can share up to five indices: protagonist , time , space , causality , and intentionality . We present the first in a planned series of evaluations that will explore increasingly nuanced methods of using these indices to predict salience. The Pairwise Event Salience Hypothesis states that when a past event shares one or more indices with the most recently narrated event, that past event is more salient than one which shares no indices with the most recently narrated event. A crowd-sourced (n=200) study of 24 short text stories that control for content, text, and length supports this hypothesis. While this is encouraging, we believe it also motivates the development of a richer model that accounts for intervening events, narrative complexity, and episodic memory decay.
LoRUS: A Mobile Crowdsourcing System for Efficiently Retrieving the Top-k Relevant Users in a Spatial Window
Mondal, Anirban (Xerox Research Center India) | Raravi, Gurulingesh (Xerox Research Center India) | Chugh, Amandeep (Xerox Research Center India) | Mukherjee, Tridib (Xerox Research Center India)
Hence, they do not address mobile resource devices, it has now become practically feasible to enable constraints (e.g., energy, bandwidth) and also result in unnecessary people to share information about dynamic events (e.g., trees spam. On the other hand, multi-cast approaches randomly fallen on roads due to a storm, sudden truck breakdowns send the queries to some of the users to preserve mobile and unscheduled processions) in their current location. This resources, but they do not ensure the direction of queries strongly motivates facilitation of various kinds of locationdependent to the most relevant users.