Country
Inferring Latent User Properties from Texts Published in Social Media
Volkova, Svitlana (Johns Hopkins University) | Bachrach, Yoram (Microsoft Research) | Armstrong, Michael ( Microsoft Research ) | Sharma, Vijay ( Microsoft Research )
We demonstrate an approach to predict latent personal attributes including user demographics, online personality, emotions and sentiments from texts published on Twitter. We rely on machine learning and natural language processing techniques to learn models from user communications. We first examine individual tweets to detect emotions and opinions emanating from them, and then analyze all the tweets published by a user to infer latent traits of that individual. We consider various user properties including age, gender, income, education, relationship status, optimism and life satisfaction. We focus on Ekmanโs six emotions: anger, joy, surprise, fear, disgust and sadness. Our work can help social network users to understand how others may perceive them based on how they communicate in social media, in addition to its evident applications in online sales and marketing, targeted advertising, large scale polling and healthcare analytics.
The Network Data Repository with Interactive Graph Analytics and Visualization
Rossi, Ryan (Purdue University) | Ahmed, Nesreen (Purdue University)
NetworkRepository (NR) is the first interactive data repository with a web-based platform for visual interactive analytics. Unlike other data repositories (e.g., UCI ML Data Repository, and SNAP), the network data repository (networkrepository.com) allows users to not only download, but to interactively analyze and visualize such data using our web-based interactive graph analytics platform. Users can in real-time analyze, visualize, compare, and explore data along many different dimensions. The aim of NR is to make it easy to discover key insights into the data extremely fast with little effort while also providing a medium for users to share data, visualizations, and insights. Other key factors that differentiate NR from the current data repositories is the number of graph datasets, their size, and variety. While other data repositories are static, they also lack a means for users to collaboratively discuss a particular dataset, corrections, or challenges with using the data for certain applications. In contrast, NR incorporates many social and collaborative aspects that facilitate scientific research, e.g., users can discuss each graph, post observations, and visualizations.
World WordNet Database Structure: An Efficient Schema for Storing Information of WordNets of the World
Redkar, Hanumant Harichandra (Indian Institute of Technology Bombay) | Bhingardive, Sudha Baban (Indian Institute of Technology Bombay) | Kanojia, Diptesh (Indian Institute of Technology Bombay) | Bhattacharyya, Pushpak (Indian Institute of Technology Bombay)
WordNet is an online lexical resource which expresses unique concepts in a language. English WordNet is the first WordNet which was developed at Princeton University. Over a period of time, many language WordNets were developed by various organizations all over the world. It has always been a challenge to store the WordNet data. Some WordNets are stored using file system and some WordNets are stored using different database models. In this paper, we present the World WordNet Database Structure which can be used to efficiently store the WordNet information of all languages of the World. This design can be adapted by most language WordNets to store information such as synset data, semantic and lexical relations, ontology details, language specific features, linguistic information, etc. An attempt is made to develop Application Programming Interfaces to manipulate the data from these databases. This database structure can help in various Natural Language Processing applications like Multilingual Information Retrieval, Word Sense Disambiguation, Machine Translation, etc.
Gene Selection in Microarray Datasets Using Progressively Refined PSO Scheme
Prasad, Yamuna (Indian Institute of Technology Delhi) | Biswas, K. K. (Indian Institute of Technology Delhi)
In this paper we propose a wrapper based PSO method for gene selection in microarray datasets, where we gradually refine the feature (gene) space from a very coarse level to a fine grained one, by reducing the gene set at each step of the algorithm. We use the linear support vector machine weight vector to serve as the initial gene pool selection. In addition, we also examine integration of other filter based ranking methods with our proposed approach. Experiments on publicly available datasets, Colon, Leukemia and T2D show that our approach selects only a very small subset of genes while yielding substantial improvements in accuracy over state-of-the-art evolutionary methods.
Salient Object Detection via Objectness Proposals
Nguyen, Tam Van (Singapore Polytechnic)
Salient object detection has gradually become a popular topic in robotics and computer vision research. This paper presents a real-time system that detects salient object by integrating objectness, foreground and compactness measures. Our algorithm consists of four basic steps. First, our method generates the objectness map via object proposals. Based on the objectness map, we estimate the background margin and compute the corresponding foreground map which prefers the foreground objects. From the objectness map and the foreground map, the compactness map is formed to favor the compact objects. We then integrate those cues to form a pixel-accurate saliency map which covers the salient objects and consistently separates fore- and background.
Visualization Techniques for Topic Model Checking
Murdock, Jaimie (Indiana University) | Allen, Colin (Indiana University)
Topic models remain a black box both for modelers and for end users in many respects. From the modelers' perspective, many decisions must be made which lack clear rationales and whose interactions are unclear โ for example, how many topics the algorithms should find (K), which words to ignore (aka the "stop list"), and whether it is adequate to run the modeling process once or multiple times, producing different results due to the algorithms that approximate the Bayesian priors. Furthermore, the results of different parameter settings are hard to analyze, summarize, and visualize, making model comparison difficult. From the end users' perspective, it is hard to understand why the models perform as they do, and information-theoretic similarity measures do not fully align with humanistic interpretation of the topics. We present the Topic Explorer, which advances the state-of-the-art in topic model visualization for document-document and topic-document relations. It brings topic models to life in a way that fosters deep understanding of both corpus and models, allowing users to generate interpretive hypotheses and to suggest further experiments. Such tools are an essential step toward assessing whether topic modeling is a suitable technique for AI and cognitive modeling applications.
Cognitive Master Teacher
Krishnapuram, Raghu (IBM Research) | Lastras, Luis A (IBM Watson Group) | Nitta, Satya (IBM Research)
The โCognitive Master Teacherโ is a result of discussions with teachers, members of educational institutions, government bodies and other thought leaders in the United States who have helped us shape its the requirements. It is conceived as a cloud-based and mobile-accessible personal agent that is readily available for teachers to use at anytime and assist them with various issues related to day-to-day teaching activities as well as professional development.
Multi-Agent Dynamic Coupling for Cooperative Vehicles Modeling
Guรฉriau, Maxime (Universitรฉ de Lyon) | Billot, Romain (Universitรฉ de Lyon) | Faouzi, Nour-Eddin El (Universitรฉ de Lyon) | Hassas, Salima (Universitรฉ de Lyon) | Armetta, Frรฉdรฉric (Universitรฉ de Lyon)
Cooperative Intelligent Transportation Systems (C-ITS) are complex systems well-suited to a multi-agent modeling. We propose a multi-agent based modeling of a C-ITS, that couples 3 dynamics (physical, informational and control dynamics) in order to ensure a smooth cooperation between non cooperative and cooperative vehicles, that communicate with each other (V2V communication) and the infrastructure (I2V and V2I communication). We present our multi-agent model, tested through simulations using real traffic data and integrated into our extension of the Multi-model Open-source Vehicular-traffic SIMulator (MovSim).
VecLP: A Realtime Video Recommendation System for Live TV Programs
Gao, Sheng (PRIS - Beijing University of Posts and Telecommunications) | Zhang, Dai (PRIS - Beijing University of Posts and Telecommunications) | Zhang, Honggang (PRIS - Beijing University of Posts and Telecommunications) | Huang, Chao (DOCOMO Beijing Communication Labs) | Zhang, Yongsheng (DOCOMO Beijing Communication Labs) | Liao, Jianxin (Beijing University of Posts and Telecommunications) | Guo, Jun (PRIS - Beijing University of Posts and Telecommunications)
We propose VecLP, a novel Internet Video recommendation system working for Live TV Programs in this paper. Given little information on the live TV programs, our proposed VecLP system can effectively collect necessary information on both the programs and the subscribers as well as a large volume of related online videos, and then recommend the relevant Internet videos to the subscribers. For that, the key frames are firstly detected from the live TV programs, and then visual and textual features are extracted from these frames to enhance the understanding of the TV broadcasts. Furthermore, by utilizing the subscribers' profiles and their social relationships, a user preference model is constructed, which greatly improves the diversity of the recommendations in our system. The subscriber's browsing history is also recorded and used to make a further personalized recommendation. This work also illustrates how our proposed VecLP system makes it happen. Finally, we dispose some sort of new recommendation strategies in use at the system to meet special needs from diverse live TV programs and throw light upon how to fuse these strategies.
CrowdMR: Integrating Crowdsourcing with MapReduce for AI-Hard Problems
Chen, Jun (Tsinghua University) | Wang, Chaokun (Tsinghua University) | Bai, Yiyuan (Tsinghua University)
Large-scale distributed computing has made available the resources necessary to solve "AI-hard" problems. As a result, it becomes feasible to automate the processing of such problems, but accuracy is not very high due to the conceptual difficulty of these problems. In this paper, we integrated crowdsourcing with MapReduce to provide a scalable innovative human-machine solution to AI-hard problems, which is called CrowdMR. In CrowdMR, the majority of problem instances are automatically processed by machine while the troublesome instances are redirected to human via crowdsourcing. The results returned from crowdsourcing are validated in the form of CAPTCHA (Completely Automated Public Turing test to Tell Computers and Humans Apart) before adding to the output. An incremental scheduling method was brought forward to combine the results from machine and human in a "pay-as-you-go" way.