Goto

Collaborating Authors

 crowdsource worker


Variational Bayesian Inference for Crowdsourcing Predictions

arXiv.org Artificial Intelligence

Crowdsourcing has emerged as an effective means for performing a number of machine learning tasks such as annotation and labelling of images and other data sets. In most early settings of crowdsourcing, the task involved classification, that is assigning one of a discrete set of labels to each task. Recently, however, more complex tasks have been attempted including asking crowdsource workers to assign continuous labels, or predictions. In essence, this involves the use of crowdsourcing for function estimation. We are motivated by this problem to drive applications such as collaborative prediction, that is, harnessing the wisdom of the crowd to predict quantities more accurately. To do so, we propose a Bayesian approach aimed specifically at alleviating overfitting, a typical impediment to accurate prediction models in practice. In particular, we develop a variational Bayesian technique for two different worker noise models - one that assumes workers' noises are independent and the other that assumes workers' noises have a latent low-rank structure. Our evaluations on synthetic and real-world datasets demonstrate that these Bayesian approaches perform significantly better than existing non-Bayesian approaches and are thus potentially useful for this class of crowdsourcing problems.


STAIR Actions: A Video Dataset of Everyday Home Actions

arXiv.org Artificial Intelligence

A new large-scale video dataset for human action recognition, called STAIR Actions is introduced. STAIR Actions contains 100 categories of action labels representing fine-grained everyday home actions so that it can be applied to research in various home tasks such as nursing, caring, and security. In STAIR Actions, each video has a single action label. Moreover, for each action category, there are around 1,000 videos that were obtained from YouTube or produced by crowdsource workers. The duration of each video is mostly five to six seconds. The total number of videos is 102,462. We explain how we constructed STAIR Actions and show the characteristics of STAIR Actions compared to existing datasets for human action recognition. Experiments with three major models for action recognition show that STAIR Actions can train large models and achieve good performance. STAIR Actions can be downloaded from http://actions.stair.center.


Crowdsourcing the Pronunciation of Out-of-Vocabulary Words

AAAI Conferences

This is an Out-of-vocabulary (OOV) words still account for a significant extremely conservative use of crowdsourcing, particularly number of the mistakes by both speech recognizers as their crowdsource workers really do speak the words in and text-to-speech synthesizers. These are not words that their experiments, rather than selecting the correct pronunciation are merely very rare, but words that were unknown to the in a multiple choice question format. Our approach lexicon used by the automatic speech recognizer (ASR) or uses nothing more than a larger number of speakers (101) text-to-speech synthesizer (TTS). In the case of ASR, even and an acoustic model in order to find the pronunciation if the pronunciation is accurately modelled, there can be a almost ab nihilo, by constructing phone lattices and submitting question as to how to spell it correctly. In the case of TTS candidate pronunciation paths to a simple weighted systems, the pronunciation of the word may be unknown, as voting algorithm that combines results across crowdsource the component euphemistically known as "letter-to-sound" workers. Our only assumption is that the basic phonetic inventory or "grapheme-to-phoneme" rules may in fact not be able to is known to the acoustic model (e.g., the pronunciation infer the pronunciation from the spelling of the word, particularly of Rodriguez selected using an English acoustic model if its provenance is unknown, or the writing system is would never trill the r's). Furthermore, whereas Rutherford more logographically constructed.