Goto

Collaborating Authors

 Information Extraction


Forward and Backward Knowledge Transfer for Sentiment Classification

arXiv.org Artificial Intelligence

This paper studies the problem of learning a sequence of sentiment classification tasks. The learned knowledge from each task is retained and used to help future or subsequent task learning. This learning paradigm is called Lifelong Learning (LL). However, existing LL methods either only transfer knowledge forward to help future learning and do not go back to improve the model of a previous task or require the training data of the previous task to retrain its model to exploit backward/reverse knowledge transfer. This paper studies reverse knowledge transfer of LL in the context of naive Bayesian (NB) classification. It aims to improve the model of a previous task by leveraging future knowledge without retraining using its training data. This is done by exploiting a key characteristic of the generative model of NB. That is, it is possible to improve the NB classifier for a task by improving its model parameters directly by using the retained knowledge from other tasks. Experimental results show that the proposed method markedly outperforms existing LL baselines.


One-shot Information Extraction from Document Images using Neuro-Deductive Program Synthesis

arXiv.org Artificial Intelligence

Our interest in this paper is in meeting a rapidly growing industrial With the rapid advancement of Deep Learning (DL) for computer demand for information extraction from images of documents such vision problems, many DL architectures are available today for as invoices, bills, receipts etc. In practice users are able to provide a document image understanding ([11], [18], [22], [28]). But like most very small number of example images labeled with the information DLbased techniques, training these models from scratch is resource that needs to be extracted. We adopt a novel'two-level''neurodeductive', and data intensive. This is a major stumbling block for industrial approach where (a) we use pre-trained deep neural problems for which collecting and annotating data incur significant networks to populate a relational database with facts about each costs in time and money. In this paper, we use two complementary document-image; and (b) we use a form of deductive reasoning, forms learning to address this problem: related to meta-interpretive learning of transition systems to learn extraction programs: Given task-specific transitions defined using (1) Neural-learning: Using pre-trained DL models for reading the entities and relations identified by the neural detectors and document images and converting them into a structured a small number of instances (usually 1, sometimes 2) of images form by populating a predefined database schema.


Gradual Machine Learning for Aspect-level Sentiment Analysis

arXiv.org Machine Learning

The state-of-the-art solutions for Aspect-Level Sentiment Analysis (ALSA) are built on a variety of deep neural networks (DNN), whose efficacy depends on large amounts of accurately labeled training data. Unfortunately, high-quality labeled training data usually require expensive manual work, and are thus not readily available in many real scenarios. In this paper, we aim to enable effective machine labeling for ALSA without the requirement for manual labeling effort. Towards this aim, we present a novel solution based on the recently proposed paradigm of gradual machine learning. It begins with some easy instances in an ALSA task, which can be automatically labeled by the machine with high accuracy, and then gradually labels the more challenging instances by iterative factor graph inference. In the process of gradual machine learning, the hard instances are gradually labeled in small stages based on the estimated evidential certainty provided by the labeled easier instances. Our extensive experiments on the benchmark datasets have shown that the performance of the proposed approach is considerably better than its unsupervised alternatives, and also highly competitive compared to the state-of-the-art supervised DNN techniques.


Information Extraction in Insurance โ€“ Claims and Underwriting Emerj

#artificialintelligence

Claims processing and underwriting are two areas of insurance that could benefit from AI-based information extraction/document search software. That said, neither are developed use-cases for AI in insurance right now. This will likely change over time as AI becomes more accessible to businesses, perhaps with autoML or a shift in the culture of innovation at older enterprises. At that point, AI use-cases in insurance will likely move from the cost-saving benefits of document search applications to more complex machine learning systems that involve document search, machine vision, and prescriptive analytics, allowing for capabilities that drive growth, such as tailor-made insurance policies.


Crowdsourcing and Validating Event-focused Emotion Corpora for German and English

arXiv.org Artificial Intelligence

Sentiment analysis has a range of corpora available across multiple languages. For emotion analysis, the situation is more limited, which hinders potential research on cross-lingual modeling and the development of predictive models for other languages. In this paper, we fill this gap for German by constructing deISEAR, a corpus designed in analogy to the well-established English ISEAR emotion dataset. Motivated by Scherer's appraisal theory, we implement a crowdsourcing experiment which consists of two steps. In step 1, participants create descriptions of emotional events for a given emotion. In step 2, five annotators assess the emotion expressed by the texts. We show that transferring an emotion classification model from the original English ISEAR to the German crowdsourced deISEAR via machine translation does not, on average, cause a performance drop.


Can a Humanoid Robot be part of the Organizational Workforce? A User Study Leveraging Sentiment Analysis

arXiv.org Artificial Intelligence

Hiring robots for the workplaces is a challenging task as robots have to cater to customer demands, follow organizational protocols and behave with social etiquette. In this study, we propose to have a humanoid social robot, Nadine, as a customer service agent in an open social work environment. The objective of this study is to analyze the effects of humanoid robots on customers at work environment, and see if it can handle social scenarios. We propose to evaluate these objectives through two modes, namely, survey questionnaire and customer feedback. We also propose a novel approach to analyze customer feedback data (text) using sentic computing methods. Specifically, we employ aspect extraction and sentiment analysis to analyze the data. From our framework, we detect sentiment associated to the aspects that mainly concerned the customers during their interaction. This allows us to understand customers expectations and current limitations of robots as employees.


Semi-Unsupervised Lifelong Learning for Sentiment Classification: Less Manual Data Annotation and More Self-Studying

arXiv.org Artificial Intelligence

Lifelong machine learning is a novel machine learning paradigm which can continually accumulate knowledge during learning. The knowledge extracting and reusing abilities enable the lifelong machine learning to solve the related problems. The traditional approaches like Na\"ive Bayes and some neural network based approaches only aim to achieve the best performance upon a single task. Unlike them, the lifelong machine learning in this paper focuses on how to accumulate knowledge during learning and leverage them for further tasks. Meanwhile, the demand for labelled data for training also is significantly decreased with the knowledge reusing. This paper suggests that the aim of the lifelong learning is to use less labelled data and computational cost to achieve the performance as well as or even better than the supervised learning.


Text Analytics Market Appears To Improve In Time by 2026 โ€“ Business Herald

#artificialintelligence

The text analytics is a process of converting unstructured text to meaningful data to understand the customer demand, product description, and market trend. The increasing demand for insights from unstructured data is one of the major drivers for growth of the text analytics market. Due to increasing stiff competition in the market, it become important for the companies to analyze the demand of their customer, and marketing strategy, which raised the demand for unstructured data. The social media industry plays a major role for providing unstructured data to the companies. Moreover, the increasing demand for predictive analytics, growing requirement for social media analytics and changing business trends are some of the key drivers which fuels the market of text analytics globally.


How Dating Apps Evolved Through Data Hub & Spoken Ep. 27

#artificialintelligence

In this episode, we talk to Nick Saretzky, Senior Director of Project Management at Tinder, about how dating apps started out with data, most recently with Tinder data. We discuss the benefits of driving change through data insights, and what user data Tinder has at its disposal. We also talk about the impact of dating apps on how people interact, and on the changing approach to modern relationships. Listen to this episode on Spotify, iTunes, and Stitcher. You can also catch up on the previous episode of the Hub & Spoken podcast, in which Jason spoke to Kerry Dawes, Director of Digital Customer Experience at The Rank Group, on the impact of data on the digital customer experience in gambling.


ArSentD-LEV: A Multi-Topic Corpus for Target-based Sentiment Analysis in Arabic Levantine Tweets

arXiv.org Machine Learning

Sentiment analysis is a highly subjective and challenging task. Its complexity further increases when applied to the Arabic language, mainly because of the large variety of dialects that are unstandardized and widely used in the Web, especially in social media. While many datasets have been released to train sentiment classifiers in Arabic, most of these datasets contain shallow annotation, only marking the sentiment of the text unit, as a word, a sentence or a document. In this paper, we present the Arabic Sentiment Twitter Dataset for the Levantine dialect (ArSenTD-LEV). Based on findings from analyzing tweets from the Levant region, we created a dataset of 4,000 tweets with the following annotations: the overall sentiment of the tweet, the target to which the sentiment was expressed, how the sentiment was expressed, and the topic of the tweet. Results confirm the importance of these annotations at improving the performance of a baseline sentiment classifier. They also confirm the gap of training in a certain domain, and testing in another domain.