Goto

Collaborating Authors

 Information Extraction


Pars-ABSA: An Aspect-based Sentiment Analysis Dataset in Persian

arXiv.org Machine Learning

Due to the increased availability of online reviews, sentiment analysis had been witnessed a booming interest from the researchers. Sentiment analysis is a computational treatment of sentiment used to extract and understand the opinions of authors. While many systems were built to predict the sentiment of a document or a sentence, many others provide the necessary detail on various aspects of the entity (i.e. aspect-based sentiment analysis). Most of the available data resources were tailored to English and the other popular European languages. Although Persian is a language with more than 110 million speakers, to the best of our knowledge, there is not any public dataset on aspect-based sentiment analysis in Persian. This paper provides a manually annotated Persian dataset, Pars-ABSA, which is verified by 3 native Persian speakers. The dataset consists of 5114 positive, 3061 negative and 1827 neutral data samples from 5602 unique reviews. Moreover, as a baseline, this paper reports the performance of some state-of-the-art aspect-based sentiment analysis methods with a focus on deep learning, on Pars-ABSA. The obtained results are impressive compared to similar English state-of-the-art.


How To Delete Your Data From Facebook Forever

TIME - Tech

Facebook has agreed to pay a record $5 billion settlement to resolve an investigation into privacy violations, the Federal Trade Commission (FTC) announced Wednesday. The company will also create an "independent privacy committee" to ensure "greater accountability at the board of directors level," an FTC press release says. But the settlement won't affect Facebook's corporate governance structure, which lets Zuckerberg hold sway over the company's actions. Facebook has promised to clean up its act when it comes to privacy matters. But the social media giant's missteps have nonetheless cost it the trust of some users.



Free Trial Signup - Gather Twitter Data DiscoverText

#artificialintelligence

Use this information to train machine-learning classifiers to recognize relevant text and social media data. Jump into data using an interactive word CloudExplorer or build a mini topic dictionary using "defined" search.


The Great Hack: the film that goes behind the scenes of the Facebook data scandal

#artificialintelligence

Cambridge Analytica may have become the byword for a scandal, but it's not entirely clear that anyone knows exactly what that scandal is. It's more like toxic word association: "Facebook", "data", "harvested", "weaponised", "Trump" and, in this country, most controversially, "Brexit". It was a media firestorm that's yet to be extinguished, a year on from whistleblower Christopher Wylie's revelations in the Observer and the New York Times about how the company acquired the personal data of tens of millions of Facebook users in order to target them in political campaigns. This week sees the release of The Great Hack, a Netflix documentary that is the first feature-length attempt to gather all the strands of the affair into some sort of narrative โ€“ though it is one contested even by those appearing in the film. "This is not about one company," Julian Wheatland, the ex-chief operating officer of Cambridge Analytica, claims at one point. "This technology is going on unabated and will continue to go on unabated.[โ€ฆ] There was always going to be a Cambridge Analytica. It just sucks to me that it's Cambridge Analytica."


Text Analytics: the convergence of Big Data and Artificial Intelligence

#artificialintelligence

The analysis of the text content in emails, blogs, tweets, forums and other forms of textual communication constitutes what we call text analytics. Text analytics is applicable to most industries: it can help analyze millions of emails; you can analyze customers-- comments and questions in forums; you can perform sentiment analysis using text analytics by measuring positive or negative perceptions of a company, brand, or product. Text Analytics has also been called text mining, and is a subcategory of the Natural Language Processing (NLP) field, which is one of the founding branches of Artificial Intelligence, back in the 1950s, when an interest in understanding text originally developed. Currently Text Analytics is often considered as the next step in Big Data analysis. Text Analytics has a number of subdivisions: Information Extraction, Named Entity Recognition, Semantic Web annotated domain--s representation, and many more.


How Bots Can Tell When the C-Suite Is Lying

#artificialintelligence

CEOs and CFOs are decidedly more nervous when fielding questions about China during earnings calls this year. What's more, they are more likely to be deceptive with their answers. "Deception associated with questions on China has skyrocketed this quarter, up about 50% from last quarter and more than double a year ago," according to a study by text analytics provider Amenity Analytics. Amenity Analytics is one of a handful of companies that are applying natural language processing (NLP), sentiment analysis and machine learning to the financial sector, evaluating earnings calls and other public meetings to unearth information of value to an investor. It is also rare technology that offers a clear path to ROI.


Multi-modal Sentiment Analysis using Deep Canonical Correlation Analysis

arXiv.org Machine Learning

This paper learns multi-modal embeddings from text, audio, and video views/modes of data in order to improve upon down-stream sentiment classification. The experimental framework also allows investigation of the relative contributions of the individual views in the final multi-modal embedding. Individual features derived from the three views are combined into a multi-modal embedding using Deep Canonical Correlation Analysis (DCCA) in two ways i) One-Step DCCA and ii) Two-Step DCCA. This paper learns text embeddings using BERT, the current state-of-the-art in text encoders. We posit that this highly optimized algorithm dominates over the contribution of other views, though each view does contribute to the final result. Classification tasks are carried out on two benchmark datasets and on a new Debate Emotion data set, and together these demonstrate that the one-Step DCCA outperforms the current state-of-the-art in learning multi-modal embeddings.


Qwant Research @DEFT 2019: Document matching and information retrieval using clinical cases

arXiv.org Machine Learning

Task 2 is a task on semantic similarity between clinical cases and discussions. For this task, we propose an approach based on language models and evaluate the impact on the results of different preprocessings and matching techniques. For task 3, we have developed an information extraction system yielding very encouraging results accuracy-wise. We have experimented two different approaches, one based on the exclusive use of neural networks, the other based on a linguistic analysis.


Wayfair Walkout, Facebook Data Value, and More News

#artificialintelligence

Tech employees are taking a stand against migrant detention centers; a proposal asking tech companies to disclose the value of your data; and a live reading of the Mueller report. Here's the news you need to know, in two minutes or less. Want to receive this two-minute roundup as an email every weekday? This afternoon, 550 employees at the Boston-based ecommerce company Wayfair staged a walkout opposing sale of company furniture to migrant detention centers. Last week, Wayfair workers discovered an order for $200,000 worth of beds and other furniture reportedly placed by government contractor BCFS for a new detention center in Carrizo Springs, Texas.