Goto

Collaborating Authors

 Information Extraction


Doing text analytics for Digital Humanities and Social Sciences with CLARIN (LDK tutorial), Galway 2017

VideoLectures.NET

Text is a basic material, a primary data layer, in many areas of humanities and social sciences. If we want to move forward with the agenda that the fields of digital humanities and computational social sciences are projecting, it is vital to bring together the technical areas that deal with automated text processing, and scholars in the humanities and social sciences. Much progress has been made in the last two decades in text analytics, a field that draws on recent advances in computational linguistics, information retrieval and machine learning. By now we know what to expect from basic tools, such as named entity recognition. To foster new areas of research, it is necessary to not only understand what is out there in terms of proven technologies and infrastructures such as CLARIN, but also how the developers of text analytics can work with researchers in the humanities and social sciences to understand the challenges in each other's field better.


Probabilistic Graphical Models for Credibility Analysis in Evolving Online Communities

arXiv.org Machine Learning

One of the major hurdles preventing the full exploitation of information from online communities is the widespread concern regarding the quality and credibility of user-contributed content. Prior works in this domain operate on a static snapshot of the community, making strong assumptions about the structure of the data (e.g., relational tables), or consider only shallow features for text classification. To address the above limitations, we propose probabilistic graphical models that can leverage the joint interplay between multiple factors in online communities --- like user interactions, community dynamics, and textual content --- to automatically assess the credibility of user-contributed online content, and the expertise of users and their evolution with user-interpretable explanation. To this end, we devise new models based on Conditional Random Fields for different settings like incorporating partial expert knowledge for semi-supervised learning, and handling discrete labels as well as numeric ratings for fine-grained analysis. This enables applications such as extracting reliable side-effects of drugs from user-contributed posts in healthforums, and identifying credible content in news communities. Online communities are dynamic, as users join and leave, adapt to evolving trends, and mature over time. To capture this dynamics, we propose generative models based on Hidden Markov Model, Latent Dirichlet Allocation, and Brownian Motion to trace the continuous evolution of user expertise and their language model over time. This allows us to identify expert users and credible content jointly over time, improving state-of-the-art recommender systems by explicitly considering the maturity of users. This also enables applications such as identifying helpful product reviews, and detecting fake and anomalous reviews with limited information.


Sentiment Analysis: Overview, Applications and Benefits

#artificialintelligence

When experimenting with machine learning and big data, you may identify data sets that contain streams of text that contain customer reviews, or social media posts where customers (or potential customers) are talking about a product, brand or service that you offer. Mining such data to determine how people feel about your product, brand, or service, is called Sentiment Analysis. People have always had an interest in what people think, or what their opinion is. Since the inception of the internet, increasing numbers of people are using websites and services to express their opinion. With social media channels such as Facebook, LinkedIn, and Twitter, it is becoming feasible to automate and gauge what public opinion is on a given topic, news story, product, or brand.


Book: Text Analytics with Python

@machinelearnbot

Text analytics can be a bit overwhelming and frustrating at times with the unstructured and noisy nature of textual data and the vast amount of information available. "Text Analytics with Python" published by Apress\Springer, is a book packed with 385 pages of useful information based on techniques, algorithms, experiences and various lessons learnt over time in analyzing text data. Learn the techniques related to natural language processing and text analytics, and gain the skills to know which technique is best suited to solve a particular problem. Text Analytics with Python teaches you both basic and advanced concepts, including text and language syntax, structure, semantics. You will focus on algorithms and techniques, such as text classification, clustering, topic modeling, and text summarization.


Text analytics: not just for customer sentiment

#artificialintelligence

Sentiment analysis is one of the most prevalent uses of text analytics, but the technology has many other valuable uses. Text analytics finds a range of applications in scientific, medical and technology development. It can detect root causes of events and augment the knowledge of what happened with an understanding of why it happened. When used predictively, it can help anticipate future outcomes and prevent adverse events. Text analytics can also enable process automation and case management.


Learn Everything about Sentiment Analysis using R

@machinelearnbot

For our case we only consider Text feature of the Tweet as we are interested on the review of the movie. We can also use the other features such as Latitude/Longitude, replied to, etc. do other analysis on the tweeted data.


Virgin Mobile makes Twitter 'free' to access

Engadget

If you have a 4G plan with Virgin Mobile, you can now access Twitter without diving in to your monthly data allowance. That means you can scroll through your feed, check your mentions and respond to pressing Direct Messages without fear of incurring any charges. The "data-free" access joins Facebook Messenger and WhatsApp, which the company first offered to subscribers last November. The only catch is that you can't stream live video through the app -- so if you want to watch the news or catch up with the day's Wimbledon action, you'll need to look elsewhere. Virgin Media says the expansion is part of a larger "plan" to offer data-free social messaging.


in-the-research-spotlight-zornitsa-kozareva

#artificialintelligence

As AWS continues to support the Artificial Intelligence (AI) community with contributions to Apache MXNet and the release of Amazon Lex, Amazon Polly, and Amazon Rekognition managed services, we are also expanding our team of AI experts, who have one primary mission: To lower the barrier to AI for all AWS developers, making AI more accessible and easy to use. At ISI, she spearheaded multimillion-dollar research grants funded by the Defense Advanced Research Projects Agency (DARPA) and Intelligence Advanced Research Projects Activity (IARPA). The research focused on topics such as machine reading, which aims at teaching machines to read and understand text just like humans do; information extraction from unstructured documents on the Web; metaphor interpretation; and sentiment analysis. Product Marketing Manager for the AWS AI portfolio of services which includes Amazon Lex, Amazon Polly, and Amazon Rekognition, as well the AWS marketing initiatives with Apache MXNet.


AI and machine learning on social media data is giving hedge funds a competitive edge

#artificialintelligence

Extracting value from a universe of data, analysing sentiment around company names (equities) or about anything else (macro), is a complex journey and we are only about 5% down that road. The parameters are evolving by which an ever-expanding data set, including the likes of Twitter, pictures, text, video is processed; relying on experts versus the wisdom of the crowd; sentiment derived from a "bag of words", as opposed to structured linguistic analysis. Last week's Unicom conference, AI, Machine Learning and Sentiment Analysis Applied to Finance (July 14) brought together a group of experts in this area. Professor Gautum Mitra, OptiRisk Systems introduced Elijah DePalma and James Cantarella, Thomson Reuters; Pierce Crosby, StockTwits; Anders Bally, Sentifi; Peter Hafez, RavenPack; Stephen Morse, Twitter. DePalma differed somewhat from the others because the Thomson Reuters sentiment engine uses only accredited Reuters news data, rather than raw social media chatter.


AI And Machine Learning On Social Media Data Is Giving Hedge Funds A Competitive Edge

International Business Times

Extracting value from a universe of data, analysing sentiment around company names (equities) or about anything else (macro), is a complex journey and we are only about 5% down that road. The parameters are evolving by which an ever-expanding data set, including the likes of Twitter, pictures, text, video is processed; relying on experts versus the wisdom of the crowd; sentiment derived from a "bag of words", as opposed to structured linguistic analysis. Last week's Unicom conference, AI, Machine Learning and Sentiment Analysis Applied to Finance (July 14) brought together a group of experts in this area. Professor Gautum Mitra, OptiRisk Systems introduced Elijah DePalma and James Cantarella, Thomson Reuters; Pierce Crosby, StockTwits; Anders Bally, Sentifi; Peter Hafez, RavenPack; Stephen Morse, Twitter. DePalma differed somewhat from the others because the Thomson Reuters sentiment engine uses only accredited Reuters news data, rather than raw social media chatter.