Information Extraction
Bluemix: Using dashDB and Insights for Twitter services to collect and store Twitter data
As part of my Technology and Innovation MBA program at Ted Rogers School of Management, I took a data and knowledge management course which teaches students the principles and practices of knowledge management. The second part of the course delves on tools used in data management and analytics. Although the theoretical part of the course was a bit dry, the hands-on portion was very interesting and exposed students to several different tools to capture, clean and analyze data. One of the tasks given to students was to capture and analyze twitter data. Although students had access to Netlytics, which is a neat cloud-based text and social network analysis tool that also collects Twitter data, students were encouraged to find other ways to collect Twitter data.
A machine-learning system that trains itself by surfing the web
MIT researchers have designed a new machine-learning system that can learn by itself to extract text information for statistical analysis when available data is scarce. This new "information extraction" system turns machine learning on its head. It works like humans do. When we run out of data in a study (say, differentiating between fake and real news), we simply search the Internet for more data, and then we piece the new data together to make sense out of it all. That differs from most machine-learning systems, which are fed as many training examples as possible to increase the chances that the system will be able to handle difficult problems by looking for patterns compared to training data.
A machine-learning system that trains itself by surfing the web
Most successful information extraction systems operate with access to a large collection of documents. In this work, we explore the task of acquiring and incorporating external evidence to improve extraction accuracy in domains where the amount of training data is scarce. This process entails issuing search queries, extraction from new sources and reconciliation of extracted values, which are repeated until sufficient evidence is collected. We approach the problem using a reinforcement learning framework where our model learns to select optimal actions based on contextual information. We employ a deep Qnetwork, trained to optimize a reward function that reflects extraction accuracy while penalizing extra effort. Our experiments on two databases – of shooting incidents, and food adulteration cases – demonstrate that our system significantly outperforms traditional extractors and a competitive meta-classifier baseline.
15 Great Blogs Posted in the last 12 Months
This is part of a new series of articles: once or twice a month, we post previous articles that were very popular when first published. These articles are at least 6 month old but no more than 12 month old. The previous digest in this series was posted here a while back. Below is our fourth edition. Top 20 Big Data Experts to Follow (Includes Scoring Algorithm) Text Classification & Sentiment Analysis tutorial / blog Learn Everything about Sentiment Analysis using R 1.5 TB dataset of anonymized user interactions released by Yahoo Fuzzy Matching Algorithms To Help Data Scientists Match Similar Data
The Emotion Journal performs real-time sentiment analysis on your most personal stories
Andrew Greenstein, an app developer from San Francisco, started journaling a few months ago. He tries to write for five minutes every day, but it's challenging to set aside the time. Still, he's read that journaling reduces stress and can help with goal-setting, so he's trying to make it a habit. At the Disrupt London Hackathon, Greenstein and his team built The Emotion Journal, a voice journaling app that performs real-time emotional analysis to detect the user's feelings and chart their emotional state over time. By day, Greenstein is the CEO of SF AppWorks, a digital agency.
DiscoverText
With dozens of powerful text mining features, including access to free and premium Gnip Twitter data, DiscoverText provides software tools to quickly and accurately evaluate text data. Data scientists know that cleaning data can be very time consuming. Users of DiscoverText build custom machine classifiers or "sifters" to find the most relevant items. DiscoverText shortens a process that used to last weeks or months; our machine-learning sifters are created in hours or just a few minutes. We support technical integrations with Twitter and SurveyMonkey.
Opinion Mining - Extraction of opinions from free text - Dataconomy
So you report with reasonable accuracies what the sentiment about a particular brand or product is. After publishing this report, your client comes back to you and says "Hey this is good. Now can you tell me ways in which I can convert the negative sentiments into positive sentiments?" – Sentiment Analysis stops there and we enter the realms of Opinion Mining. Opinion Mining is about having a deeper understanding of the review that was written. Typically, a detailed review will not just have a sentiment attached to it. It will have information and valuable feedback that can literally help to build the next strategy.
Text Analytics and Machine Learning: A Virtuous Combination
The world of big data analytics is incredibly diverse, and people are coming up with new analytic tools and techniques every day. But one particularly productive combination that should not be overlooked involves the use of text analytics and machine learning. Tom Sabo, principal solutions architect at analytics giant SAS, says the one-two punch of predictive modeling on structured data, and text mining with unstructured data, can deliver insights that are more than the sum of their analytic parts. "They really run side by side," Sabo tells Datanami. "Let's say somebody has predictive models in place against whether customer will churn or to maximize profit, for instance. If they have text, like notes, in the rest of that structured data…we can incorporate that additional free form information for actionable insight."
Smart Business: automated sentiments analysis on top
The modern world seems really fast and dynamic with a multitude of new products being launched. Marketing agencies are making fortune by monitoring the markets and delivering reports on consumers' opinions. For today, the feedback analysis is a separate area, let's say a growing industry with an array of products and services. And the prices for those services are pretty exorbitant. Without any doubts, there's always an opportunity to start personal volcanic activities on feedback collection and analysis.