Goto

Collaborating Authors

 Information Extraction


Twitter discussions and emotions about COVID-19 pandemic: a machine learning approach

arXiv.org Machine Learning

The objective of the study is to examine coronavirus disease (COVID-19) related discussions, concerns, and sentiments that emerged from tweets posted by Twitter users. We analyze 4 million Twitter messages related to the COVID-19 pandemic using a list of 25 hashtags such as "coronavirus," "COVID-19," "quarantine" from March 1 to April 21 in 2020. We use a machine learning approach, Latent Dirichlet Allocation (LDA), to identify popular unigram, bigrams, salient topics and themes, and sentiments in the collected Tweets. Popular unigrams include "virus," "lockdown," and "quarantine." Popular bigrams include "COVID-19," "stay home," "corona virus," "social distancing," and "new cases." We identify 13 discussion topics and categorize them into five different themes, such as "public health measures to slow the spread of COVID-19," "social stigma associated with COVID-19," "coronavirus news cases and deaths," "COVID-19 in the United States," and "coronavirus cases in the rest of the world". Across all identified topics, the dominant sentiments for the spread of coronavirus are anticipation that measures that can be taken, followed by a mixed feeling of trust, anger, and fear for different topics. The public reveals a significant feeling of fear when they discuss the coronavirus new cases and deaths than other topics. The study shows that Twitter data and machine learning approaches can be leveraged for infodemiology study by studying the evolving public discussions and sentiments during the COVID-19. Real-time monitoring and assessment of the Twitter discussion and concerns can be promising for public health emergency responses and planning. Already emerged pandemic fear, stigma, and mental health concerns may continue to influence public trust when there occurs a second wave of COVID-19 or a new surge of the imminent pandemic.


Comparative Sentiment Analysis of App Reviews

arXiv.org Machine Learning

Google app market captures the school of thought of users via ratings and text reviews. The critique's viewpoint regarding an app is proportional to their satisfaction level. Consequently, this helps other users to gain insights before downloading or purchasing the apps. The potential information from the reviews can't be extracted manually, due to its exponential growth. Sentiment analysis, by machine learning algorithms employing NLP, is used to explicitly uncover and interpret the emotions. This study aims to perform the sentiment classification of the app reviews and identify the university students' behavior towards the app market. We applied machine learning algorithms using the TF-IDF text representation scheme and the performance was evaluated on the ensemble learning method. Our model was trained on Google reviews and tested on students' reviews. SVM recorded the maximum accuracy(93.37\%), F-score(0.88) on tri-gram + TF-IDF scheme. Bagging enhanced the performance of LR and NB with accuracy of 87.80\% and 85.5\% respectively.


Leveraging Multimodal Behavioral Analytics for Automated Job Interview Performance Assessment and Feedback

arXiv.org Machine Learning

Behavioral cues play a significant part in human communication and cognitive perception. In most professional domains, employee recruitment policies are framed such that both professional skills and personality traits are adequately assessed. Hiring interviews are structured to evaluate expansively a potential employee's suitability for the position - their professional qualifications, interpersonal skills, ability to perform in critical and stressful situations, in the presence of time and resource constraints, etc. Therefore, candidates need to be aware of their positive and negative attributes and be mindful of behavioral cues that might have adverse effects on their success. We propose a multimodal analytical framework that analyzes the candidate in an interview scenario and provides feedback for predefined labels such as engagement, speaking rate, eye contact, etc. We perform a comprehensive analysis that includes the interviewee's facial expressions, speech, and prosodic information, using the video, audio, and text transcripts obtained from the recorded interview. We use these multimodal data sources to construct a composite representation, which is used for training machine learning classifiers to predict the class labels. Such analysis is then used to provide constructive feedback to the interviewee for their behavioral cues and body language. Experimental validation showed that the proposed methodology achieved promising results.


Weakly-supervised Domain Adaption for Aspect Extraction via Multi-level Interaction Transfer

arXiv.org Artificial Intelligence

Fine-grained aspect extraction is an essential sub-task in aspect based opinion analysis. It aims to identify the aspect terms (a.k.a. opinion targets) of a product or service in each sentence. However, expensive annotation process is usually involved to acquire sufficient token-level labels for each domain. To address this limitation, some previous works propose domain adaptation strategies to transfer knowledge from a sufficiently labeled source domain to unlabeled target domains. But due to both the difficulty of fine-grained prediction problems and the large domain gap between domains, the performance remains unsatisfactory. This work conducts a pioneer study on leveraging sentence-level aspect category labels that can be usually available in commercial services like review sites to promote token-level transfer for the extraction purpose. Specifically, the aspect category information is used to construct pivot knowledge for transfer with assumption that the interactions between sentence-level aspect category and token-level aspect terms are invariant across domains. To this end, we propose a novel multi-level reconstruction mechanism that aligns both the fine-grained and coarse-grained information in multiple levels of abstractions. Comprehensive experiments demonstrate that our approach can fully utilize sentence-level aspect category labels to improve cross-domain aspect extraction with a large performance gain.


Part-1: Introduction to Natural Language Processing (NLP)

#artificialintelligence

Natural language processing (NLP) is a field of artificial intelligence in which computers analyze, understand, and derive meaning information from human language in a smart and useful way. By utilizing NLP, developers can organize and structure knowledge to perform tasks such as automatic summarization, translation, named entity recognition, relationship extraction, sentiment analysis, speech recognition, and topic segmentation. NLP is characterized as a difficult problem in computer science. Human language is rarely precise or plainly spoken. To understand human language is to understand not only the words but the concepts and how they're linked together to create meaning.


Social Sentiment Analysis Toward the Clean Energy Transition

#artificialintelligence

The world is in the midst of an energy transition. This massive shift aims to move away from reliance on fuels that are destructive to the climate, the environment, and people's well-being. The goal established by the UN is to "ensure access to affordable, reliable, sustainable and modern energy for all" by 2030. While governments, energy companies, and activists dominate the headlines, the progress with infrastructure and technology won't be sufficient. A successful energy transition for the good of all humanity depends on the action of individuals.


What is NLP and Why is it Important?

#artificialintelligence

Natural Language Processing (NLP) is a subfield of artificial intelligence that assists computers with understanding human language. Utilizing NLP, machines can understand unstructured online information so we can gain significant insights. As computer technology advances past their artificial requirements, companies are searching for better approaches to exploit. A sharp increase in computing speed and capacities has led to new and highly intelligent software systems, some of which are prepared to supplant or augment human services. The rise of natural language processing (NLP) is probably the best example, with intelligent chatbots prepared to change the universe of customer service and beyond.


GeoCoV19: A Dataset of Hundreds of Millions of Multilingual COVID-19 Tweets with Location Information

arXiv.org Artificial Intelligence

The past several years have witnessed a huge surge in the use of social media platforms during mass convergence events such as health emergencies, natural or human-induced disasters. These non-traditional data sources are becoming vital for disease forecasts and surveillance when preparing for epidemic and pandemic outbreaks. In this paper, we present GeoCoV19, a large-scale Twitter dataset containing more than 524 million multilingual tweets posted over a period of 90 days since February 1, 2020. Moreover, we employ a gazetteer-based approach to infer the geolocation of tweets. We postulate that this large-scale, multilingual, geolocated social media data can empower the research communities to evaluate how societies are collectively coping with this unprecedented global crisis as well as to develop computational methods to address challenges such as identifying fake news, understanding communities' knowledge gaps, building disease forecast and surveillance models, among others.


New Cognitive Services capabilities are now generally available Azure updates Microsoft Azure

#artificialintelligence

Computer Vision--Advanced text extraction: The most advanced text extraction capability for Computer Vision, Read 3.0, is now generally available and expanding its language coverage beyond English and Spanish to include French, German, Portuguese, Italian, and Dutch. Read 3.0 in containers is also available in preview. Language Understanding and Text Analytics sentiment analysis in containers are now generally available. Language Understanding--Enhanced portal experience: The Language Understanding service has revamped the labeling experience, making it easier to build apps and bots that can understand the complex language structures people tend to use. For example, in this order: "I want a large chicken pizza without sauce and a medium pizza with olives," there are two different language structures within the same order.


How to Build Emotion Text Analyzer with Python (NLP)

#artificialintelligence

In this tutorial I will guide you on how to detect emotions associated with textual data which can be either classified as either positive or negative and how can you apply that knowledge in variety of applications depending on what you wanna do. In this tutorial I will guide you on how to detect emotions associated with textual data which can be either classified as either positive or negative and how can you apply that knowledge in variety of applications depending on what you wanna do. For instance you want to perform automatic analysis of customer feedback with directly reading them as either positive or negative feedback you will need to Sentiment analyzer to check the negativity or positivity of the textual data.