Asia
A Markov Decision Process Framework for Predictable Job Completion Times on Crowdsourcing Platforms
Lakshminarayanan, Chandrashekar (Indian Institute of Science) | Dubey, Ayush (Indian Institute of Science) | Bhatnagar, Shalabh (Indian Institute of Science) | Balamurugan, Chithralekha (Xerox Research Centre India)
Task starvation leads to huge variation in the completion times of the tasks posted on to the crowd. The price offered to a given task together with the dynamics of the crowd at the time of posting affect its completion time. Large organizations/requesters who frequent the crowd at regular intervals in order to get their tasks done desire predictability in completion times of the tasks. Thus, such requesters have to take into account the crowd dynamics at the time of posting the tasks and price them accordingly. In this work, we study an instance of the pricing problem and propose a solution based on the framework of Markov Decision Processes (MDPs).
Crowd-Training Machine Learning Systems for Human Rights Abuse Documentation
Aronson, Jay D. (Carnegie Mellon University)
In this talk, I will describe efforts being undertaken in a collaboration between human rights advocates and Social media and mobile phones with good cameras and computer scientists at Carnegie Mellon University to Internet access are dramatically changing the nature of develop tools, methods and algorithms that will make it human rights documentation, reporting and advocacy. Key to this process, and like YouTube, Live Leak, Vimeo, and Facebook every apropos of this session, is the development of mechanisms week. In Syria, more than 650,000 videos have been to enable "the crowd" (i.e., those individuals around the uploaded to social media sites since the conflict started world who care about human rights and have relevant three years ago. This trove of interest dies down or moves on to new issues or places. In presenting this relevant in the long-term, what is irrelevant to the project, I hope to get feedback from other participants in situation or repetitive, and what is patently false or the workshop on how to achieve this goal, particularly by misleading.
Crowdsourced Explanations for Humorous Internet Memes Based on Linguistic Theories
Lin, Chi-Chin (National Taiwan University) | Huang, Yi-Ching (National Taiwan University) | Hsu, Jane Yung-jen (National Taiwan University)
Humorous images can be seen in many social media websites. However, newcomers to these websites often have trouble fitting in because the community subculture is usually implicit. Among all the types of humorous images, Internet memes are relatively hard for newcomers to understand. In this work, we develop a system that leverages crowdsourcing techniques to generate explanations for memes. We claim that people who are not familiar with Internet meme subculture can still quickly pick up the gist of the memes by reading the explanations. Our template-based explanations illustrate the incongruity between normal situations and the punchlines in jokes. The explanations can be produced by completing the two proposed human task processes. Experimental results suggest that the explanations produced by our system greatly help newcomers to understand unfamiliar memes. For further research, it is possible to employ our explanation generation system to improve computational humanities.
Learning Pronunciation and Accent from The Crowd
Liu, Frederick (National Taiwan University) | Yang, Jeremy Chiaming (National Taiwan University) | Hsu, Jane Yung-jen (National Taiwan University)
Learning a second language is becoming a more popular trend around the world. But the act of learning another language in a place removed from native speakers is difficult as there is often no one to correct mistakes nor examples to imitate. With the idea of crowd sourcing, we would like to propose an efficient way to learn a second language better.
Crowdsourced Data Analytics: A Case Study of a Predictive Modeling Competition
Baba, Yukino (National Institute of Informatics) | Nori, Nozomi (Kyoto University) | Saito, Shigeru (OPT, Inc.) | Kashima, Hisashi (Kyoto University)
Predictive modeling competitions provide a new data mining approach that leverages crowds of data scientists to examine a wide variety of predictive models and build the best performance model. In this paper, we report the results of a study conducted on CrowdSolving, a platform for predictive modeling competitions in Japan. We hosted a competition on a link prediction task and observed that (i) the prediction performance of the winner significantly outperformed that of a state-of-the-art method, (ii) the aggregated model constructed from all submitted models further improved the final performance, and (iii) the performance of the aggregated model built only from early submissions nevertheless overtook the final performance of the winner.
CrowdUtility: A Recommendation System for Crowdsourcing Platforms
Chander, Deepthi (Xerox Research Center India) | Bhattacharya, Sakyajit (Xerox Research Centre India) | Celis, Elisa (EPFL Lausanne) | Dasgupta, Koustuv (Xerox Research Centre India) | Karanam, Saraschandra (Xerox Research Centre India) | Rajan, Vaibhav (Xerox Research Centre India) | Gupta, Avantika (Xerox Research Centre India)
Crowd workers exhibit varying work patterns, expertise, and quality leading to wide variability in the performance of crowdsourcing platforms. The onus of choosing a suitable platform to post tasks is mostly with the requester, often leading to poor guarantees and unmet requirements due to the dynamism in performance of crowd platforms. Towards this end, we demonstrate CrowdUtility, a statistical modelling based tool for evaluating multiple crowdsourcing platforms and recommending a platform that best suits the requirements of the requester. CrowdUtility uses an online Multi-Armed Bandit framework, to schedule tasks while optimizing platform performance. We demonstrate an end-to end system starting from requirements specification, to platform recommendation, to real-time monitoring.
Predicting Own Action: Self-Fulfilling Prophecy Induced by Proper Scoring Rules
Oka, Masaaki (Kyushu University) | Todo, Taiki (Kyushu University) | Sakurai, Yuko (Kyushu University) | Yokoo, Makoto (Kyushu University)
This paper studies a mechanism to incentivize agents who predict their own future actions and truthfully declare their predictions. In a crowdsouring setting (e.g., participatory sensing), obtaining an accurate prediction of the actions of workers/agents is valuable for a requester who is collecting real-world information from the crowd. If an agent predicts an external event that she cannot control herself (e.g., tomorrow's weather), any proper scoring rule can give an accurate incentive. In our problem setting, an agent needs to predict her own action (e.g., what time tomorrow she will take a photo of a specific place) that she can control to maximize her utility. Also, her (gross) utility can vary based on an eternal event. We first prove that a mechanism can satisfy our goal if and only if it utilizes a strictly proper scoring rule, assuming that an agent can find an optimal declaration that maximizes her expected utility. This declaration is self-fulfilling; if she acts to maximize her utility, the probabilistic distribution of her action matches her declaration, assuming her prediction about the external event is correct. Furthermore, we develop a heuristic algorithm that efficiently finds a semi-optimal declaration, and show that this declaration is still self-fulfilling. We also examine our heuristic algorithm's performance and describe how an agent acts when she faces an unexpected scenario.
Quality Control for Crowdsourced Enumeration Tasks
Kajimura, Shunsuke (The University of Tokyo) | Baba, Yukino (National Institute of Informatics) | Kajino, Hiroshi (The University of Tokyo) | Kashima, Hisashi (Kyoto University)
Quality control is one of the central issues in crowdsourcing research. In this paper, we consider a quality control problem of crowdsourced enumeration tasks that request workers to enumerate possible answers as many as possible. Since workers neither necessarily provide correct answers nor provide exactly the same answers even if the answers indicate the same idea, we propose a two-stage quality control method consisting of the answer clustering stage and the reliability estimation stage.
Speech Synthesis Data Collection for Visually Impaired Person
Ashikawa, Masayuki (Toshiba Corporation) | Kawamura, Takahiro (Toshiba Corporation) | Ohsuga, Akihiko (The University of Electro-Communications)
Crowdsourcing platforms provide attractive solutions for collecting speech synthesis data for visually impaired person. However, quality control problems remain because of low-quality volunteer workers. In this paper, we propose the design of a crowdsourcing system that allows us to devise quality control methods. We introduce four worker selection methods; preprocessing filtering, real-time filtering, post-processing filtering, and guess-processing filtering. These methods include a novel approach that utilizes a collaborative filtering technique in addition to a basic approach involving initial training or use of gold-standard data. These quality control methods improved the quality of collected speech synthesis data. Moreover, we have already collected 140,000 Japanese words from 500 million web data for speech synthesis data.
Partition-wise Linear Models
Oiwa, Hidekazu, Fujimaki, Ryohei
Region-specific linear models are widely used in practical applications because of their non-linear but highly interpretable model representations. One of the key challenges in their use is non-convexity in simultaneous optimization of regions and region-specific models. This paper proposes novel convex region-specific linear models, which we refer to as partition-wise linear models. Our key ideas are 1) assigning linear models not to regions but to partitions (region-specifiers) and representing region-specific linear models by linear combinations of partition-specific models, and 2) optimizing regions via partition selection from a large number of given partition candidates by means of convex structured regularizations. In addition to providing initialization-free globally-optimal solutions, our convex formulation makes it possible to derive a generalization bound and to use such advanced optimization techniques as proximal methods and decomposition of the proximal maps for sparsity-inducing regularizations. Experimental results demonstrate that our partition-wise linear models perform better than or are at least competitive with state-of-the-art region-specific or locally linear models.