Information Extraction
Microsummarization of Online Reviews: An Experimental Study
Mason, Rebecca (Google, Inc.) | Gaska, Benjamin (University of Arizona) | Durme, Benjamin Van (Johns Hopkins University) | Choudhury, Pallavi (Microsoft Research) | Hart, Ted (Microsoft Research) | Dolan, Bill (Microsoft Research) | Toutanova, Kristina (Microsoft Research) | Mitchell, Margaret (Microsoft Research)
Mobile and location-based social media applications provide platforms for users to share brief opinions about products, venues, and services. These quickly typed opinions, or microreviews, are a valuable source of current sentiment on a wide variety of subjects. However, there is currently little research on how to mine this information to present it back to users in easily consumable way. In this paper, we introduce the task of microsummarization, which combines sentiment analysis, summarization, and entity recognition in order to surface key content to users. We explore unsupervised and supervised methods for this task, and find we can reliably extract relevant entities and the sentiment targeted towards them using crowdsourced labels as supervision. In an end-to-end evaluation, we find our best-performing system is vastly preferred by judges over a traditional extractive summarization approach. This work motivates an entirely new approach to summarization, incorporating both sentiment analysis and item extraction for modernized, at-a-glance presentation of public opinion.
Acquiring Knowledge of Affective Events from Blogs Using Label Propagation
Ding, Haibo (University of Utah) | Riloff, Ellen (University of Utah)
Many common events in our daily life affect us in positive and negative ways. For example, going on vacation is typically an enjoyable event, while being rushed to the hospital is an undesirable event. In narrative stories and personal conversations, recognizing that some events have a strong affective polarity is essential to understand the discourse and the emotional states of the affected people. However, current NLP systems mainly depend on sentiment analysis tools, which fail to recognize many events that are implicitly affective based on human knowledge about the event itself and cultural norms. Our goal is to automatically acquire knowledge of stereotypically positive and negative events from personal blogs. Our research creates an event context graph from a large collection of blog posts and uses a sentiment classifier and semi-supervised label propagation algorithm to discover affective events. We explore several graph configurations that propagate affective polarity across edges using local context, discourse proximity, and event-event co-occurrence. We then harvest highly affective events from the graph and evaluate the agreement of the polarities with human judgements.
Short Text Representation for Detecting Churn in Microblogs
Amiri, Hadi (University of Maryland) | III, Hal Daume (University of Maryland)
Churn happens when a customer leaves a brand or stop using its services. Brands reduce their churn rates by identifying and retaining potential churners through customer retention campaigns. In this paper, we consider the problem of classifying micro-posts as churny or non-churny with respect to a given brand. Motivated by the recent success of recurrent neural networks (RNNs) in word representation, we propose to utilize RNNs to learn micro-post and churn indicator representations. We show that such representations improve the performance of churn detection in microblogs and lead to more accurate ranking of churny contents. Furthermore, in this researchwe show that state-of-the-art sentiment analysis approaches fail to identify churny contents. Experiments on Twitter data about three telco brands show the utility of our approach for this task.
Building a Large Scale Dataset for Image Emotion Recognition: The Fine Print and The Benchmark
You, Quanzeng (University of Rochester) | Luo, Jiebo (University of Rochester) | Jin, Hailin (Adobe Research ) | Yang, Jianchao (Snapchat Inc)
Psychological research results have confirmed that people can have different emotional reactions to different visual stimuli. Several papers have been published on the problem of visual emotion analysis. In particular, attempts have been made to analyze and predict people's emotional reaction towards images. To this end, different kinds of hand-tuned features are proposed. The results reported on several carefully selected and labeled small image data sets have confirmed the promise of such features. While the recent successes of many computer vision related tasks are due to the adoption of Convolutional Neural Networks (CNNs), visual emotion analysis has not achieved the same level of success. This may be primarily due to the unavailability of confidently labeled and relatively large image data sets for visual emotion analysis. In this work, we introduce a new data set, which started from 3+ million weakly labeled images of different emotions and ended up 30 times as large as the current largest publicly available visual emotion data set. We hope that this data set encourages further research on visual emotion analysis. We also perform extensive benchmarking analyses on this large data set using the state of the art methods including CNNs.
Context-Sensitive Twitter Sentiment Classification Using Neural Network
Ren, Yafeng (Wuhan University) | Zhang, Yue (Singapore University of Technology and Design) | Zhang, Meishan (Heilongjiang University) | Ji, Donghong (Wuhan University)
Sentiment classification on Twitter has attracted increasing research in recent years.Most existing work focuses on feature engineering according to the tweet content itself.In this paper, we propose a context-based neural network model for Twitter sentiment analysis, incorporating contextualized features from relevant Tweets into the model in the form of word embedding vectors.Experiments on both balanced and unbalanced datasets show that our proposed models outperform the current state-of-the-art.
Identifying Sentiment Words Using an Optimization Model with L1 Regularization
Deng, Zhi-Hong (Peking University) | Yu, Hongliang (Carnegie Mellon University) | Yang, Yunlun (Peking University)
Sentiment word identification is a fundamental work in numerous applications of sentiment analysis and opinion mining, such as review mining, opinion holder finding, and twitter classification. In this paper, we propose an optimization model with L1 regularization, called ISOMER, for identifying the sentiment words from the corpus. Our model can employ both seed words and documents with sentiment labels, different from most existing researches adopting seed words only. The L1 penalty in the objective function yields a sparse solution since most candidate words have no sentiment. The experiments on the real datasets show that ISOMER outperforms the classic approaches, and that the lexicon learned by ISOMER can be effectively adapted to document-level sentiment analysis.
These U.S. States Like Bacon the Most, Based on Instagram Data
Here's a point that's difficult to argue: Bacon is delicious. The fatty pig product is a favorite food for breakfast or otherwise among many (looking at you, hipsters). Housewares retailer Ginny's recently released an interactive map dubbed the "50 States of Bacon," which uses Instagram data to show which states enjoy bacon the most – and least. According to the map, based in part on an analysis of more than 33,000 Instagram photos in the U.S. featuring the tag #bacon, Hawaii loves it the least, while Nebraskans are the most bacon-obsessed. See your state's amount of bacon love by clicking here.
Work Smarter Not Harder - Teradata Text Analytics
Text analytics is a genre of analytic capabilities intended to function across the typed/written word. This area of analytics seeks to learn from huge quantities of text data to expose human intent, sentiment, and behaviors. Examples include doctor notes, tweets, product/content reviews, survey text response, and much more. There are varying types of text analytics such as text parsing, Levenshtein distance, entity extraction, tagging/classification, and chunking - just to name a few. Many areas of machine learning use text analytics as data preparation steps in order to develop models.
Symbiotic Cognitive Computing through Iteratively Supervised Lexicon Induction
Alba, Alfredo (IBM Research) | Drews, Clemens (IBM Research) | Gruhl, Daniel (IBM Research) | Lewis, Neal (IBM Research) | Mendes, Pablo N. (IBM Research) | Nagarajan, Meenakshi (IBM Research) | Welch, Steve (IBM Research) | Coden, Anni (IBM Research) | Qadir, Ashequl (University of Utah)
In this paper we approach a subset of semantic analysis tasks through a symbiotic cognitive computing approach -- the user and the system learn from each other and accomplish the tasks better than they would do on their own. Our approach starts with a domain expert building a simplified domain model (e.g. semantic lexicons) and annotating documents with that model. The system helps the user by allowing them to obtain quicker results, and by leading them to refine their understanding of the domain. Meanwhile, through the feedback from the user, the system adapts more quickly and produces more accurate results. We believe this virtuous cycle is key for building next generation high quality semantic analysis systems. We present some preliminary findings and discuss our results on four aspects of this virtuous cycle, namely: the intrinsic incompleteness of semantic models, the need for a human in the loop, the benefits of a computer in the loop and finally the overall improvements offered by the human-computer interaction in the process.
Creating a Mars Target Encyclopedia by Extracting Information from the Planetary Science Literature
Wagstaff, Kiri L. (Jet Propulsion Laboratory) | Riloff, Ellen (University of Utah) | Lanza, Nina L. (Los Alamos National Laboratory) | Mattmann, Chris A. (Jet Propulsion Laboratory) | Ramirez, Paul M. (Jet Propulsion Laboratory)
Staying up to date with the latest discoveries is a challenge in any scientific field. In planetary science, new observation targets on the surface of Mars are identified and named every day, and new publications announcing new discoveries and conclusions provide frequent updates about these targets. We are constructing a system that uses information extraction and retrieval methods to mine the steadily growing body of planetary science publications about Mars surface targets and automatically construct a concise summary of what is known about each target. The Mars Target Encyclopedia will provide a central, continually updated resource for use by planetary scientists and the interested public. We describe our use of Tika, Sundance, and AutoSlog to extract and summarize information, some of the challenges associated with this domain, and our plans for maturing the system.