Information Extraction
Sentiment Analysis of Movie Reviews (3): doc2vec
This is the last – for now – installment of my mini-series on sentiment analysis of the Stanford collection of IMDB reviews (originally published on recurrentnull.wordpress.com). So far, we've had a look at classical bag-of-words models and word vectors (word2vec). We saw that from the classifiers used, logistic regression performed best, be it in combination with bag-of-words or word2vec. We also saw that while the word2vec model did in fact model semantic dimensions, it was less successful for classification than bag-of-words, and we explained that by the averaging of word vectors we had to perform to obtain input features on review (not word) level. So the question now is: How would distributed representations perform if we did not have to throw away information by averaging word vectors?
Opinion Mining - Sentiment Analysis and Beyond
So you report with reasonable accuracies what the sentiment about a particular brand or product is. After publishing this report, your client comes back to you and says "Hey this is good. Now can you tell me ways in which I can convert the negative sentiments into positive sentiments?" – Sentiment Analysis stops there and we enter the realms of Opinion Mining. Opinion Mining is about having a deeper understanding of the review that was written. Typically, a detailed review will not just have a sentiment attached to it. It will have information and valuable feedback that can literally help to build the next strategy.
Sentiment Analysis of Movie Reviews (1):Bag-of-Words Models
Imagine I show you a book review, on amazon.com, Imagine I hide the number of stars, – all you get to see is the number of stars. And now I'm asking you, that review, is it good or bad? Well, it should be easy, for humans (although depending on the input there can be lots of disagreement between humans, too.) But if you want to do it automatically, it turns out to be surprisingly difficult.
Sentiment Analysis of Movie Reviews (2): word2vec
This is the continuation of my mini-series on sentiment analysis of movie reviews, which originally appeared on recurrentnull.wordpress.com. Last time, we had a look at how well classical bag-of-words models worked for classification of the Stanford collection of IMDB reviews. As it turned out, the "winner" was Logistic Regression, using both unigrams and bigrams for classification. The best classification accuracy obtained was .89 So, bag-of-words models may be surprisingly successful, but they are limited in what they can do.
A machine-learning system that trains itself by surfing the web
Most successful information extraction systems operate with access to a large collection of documents. In this work, we explore the task of acquiring and incorporating external evidence to improve extraction accuracy in domains where the amount of training data is scarce. This process entails issuing search queries, extraction from new sources and reconciliation of extracted values, which are repeated until sufficient evidence is collected. We approach the problem using a reinforcement learning framework where our model learns to select optimal actions based on contextual information. We employ a deep Qnetwork, trained to optimize a reward function that reflects extraction accuracy while penalizing extra effort. Our experiments on two databases – of shooting incidents, and food adulteration cases – demonstrate that our system significantly outperforms traditional extractors and a competitive meta-classifier baseline.
A Simple Approach to Multilingual Polarity Classification in Twitter
Tellez, Eric S., Jiménez, Sabino Miranda, Graff, Mario, Moctezuma, Daniela, Suárez, Ranyart R., Siordia, Oscar S.
Recently, sentiment analysis has received a lot of attention due to the interest in mining opinions of social media users. Sentiment analysis consists in determining the polarity of a given text, i.e., its degree of positiveness or negativeness. Traditionally, Sentiment Analysis algorithms have been tailored to a specific language given the complexity of having a number of lexical variations and errors introduced by the people generating content. In this contribution, our aim is to provide a simple to implement and easy to use multilingual framework, that can serve as a baseline for sentiment analysis contests, and as starting point to build new sentiment analysis systems. We compare our approach in eight different languages, three of them have important international contests, namely, SemEval (English), TASS (Spanish), and SENTIPOLC (Italian). Within the competitions our approach reaches from medium to high positions in the rankings; whereas in the remaining languages our approach outperforms the reported results.
How to create a Twitter Sentiment Analysis using R and Shiny
I will show you how to create a simple application in R & Shiny to perform Twitter Sentiment Analysis in real-time. First, I create a Shiny Project. Then, in the ui.R file, I put this code: Here, I will show a title, the current time, a table with Twitter user name, a bar graph and wordclouds. It's a really simply code, not complex at all. The purpose of it is just for testing and so you guys can practice R language.
Impactful text analytics for smarter businesses
However, most importantly, the restaurant owner has the most scope for extracting valuable snippets of insights from customer reviews with ratings between 3-4/5. I recently had a chance to deliver a talk in a conference titled'Understanding Consumers in the Digital World', held at IIM Lucknow, Noida Campus on 16-17th November 2015. The audience mainly comprised of marketers, market research professionals and academics whose work is primarily focused on obtaining deep insights by understanding the online consumers. My talk was titled'Decoding Ratings for superior service in restaurants – Using text to understand customers'. The focus was quite simple – convince and demonstrate how to read and understand customers from their reviews, not ratings. Our product, Lunchbox, a complete restaurant management solution was showcased as well.
SuperBowl XLIX in Tweets: Sentiment Analysis of 4 Million Tweets
This blog was originally published on our Text Analysis blog, the blog post set out to analyze and visualize 4 million tweets collected during Superbowl XLIX. Not surprisingly, Superbowl XLIX generated a huge amount of chatter on social networks with Twitter Estimating that over 28.4 million posts made with terms relating to the Superbowl. At AYLIEN, we collected just under 4 million Tweets from Hashtags, Handles and Keywords we were monitoring. To keep our sample clean, we removed any reTweets and spam from the Tweets collected and only worked with those Tweets that were written in English. We were left with about 3.5 million Tweets to play with.