Goto

Collaborating Authors

 Information Extraction


Quootstrap: Scalable Unsupervised Extraction of Quotation-Speaker Pairs from Large News Corpora via Bootstrapping

AAAI Conferences

We propose Quootstrap, a method for extracting quotations, as well as the names of the speakers who uttered them, from large news corpora. Whereas prior work has addressed this problem primarily with supervised machine learning, our approach follows a fully unsupervised bootstrapping paradigm. It leverages the redundancy present in large news corpora, more precisely, the fact that the same quotation often appears across multiple news articles in slightly different contexts. Starting from a few seed patterns, such as ["Q", said S.], our method extracts a set of quotation-speaker pairs (Q, S), which are in turn used for discovering new patterns expressing the same quotations; the process is then repeated with the larger pattern set. Our algorithm is highly scalable, which we demonstrate by running it on the large ICWSM 2011 Spinn3r corpus. Validating our results against a crowdsourced ground truth, we obtain 90% precision at 40% recall using a single seed pattern, with significantly higher recall values for more frequently reported (and thus likely more interesting) quotations. Finally, we showcase the usefulness of our algorithm's output for computational social science by analyzing the sentiment expressed in our extracted quotations.


Opinion Context Extraction for Aspect Sentiment Analysis

AAAI Conferences

Sentiment analysis is the computational study of opinionated text and is becoming increasing important to online commercial applications. However, the majority of current approaches determine sentiment by attempting to detect the overall polarity of a sentence, paragraph, or text window, but without any knowledge about the entities mentioned (e.g. restaurant) and their aspects (e.g. price). Aspect-level sentiment analysis of customer feedback data when done accurately can be leveraged to understand strong and weak performance points of businesses and services, and can also support the formulation of critical action steps to improve performance. In this paper we focus on aspect-level sentiment classification, studying the role of opinion context extraction for a given aspect and the extent to which traditional and neural sentiment classifiers benefit when trained using the opinion context text. We propose four methods to aspect context extraction using lexical, syntactic and sentiment co-occurrence knowledge. Further, we evaluate the usefulness of the opinion contexts for aspect-sentiment analysis. Our experiments on benchmark data sets from SemEval and a real-world dataset from the insurance domain suggests that extracting the right opinion context is effective in improving classification performance.Specifically combining syntactical features with sentiment co-occurrence knowledge leads to the best aspect-sentiment classification performance.


Using J-K fold Cross Validation to Reduce Variance When Tuning NLP Models

arXiv.org Machine Learning

K-fold cross validation (CV) is a popular method for estimating the true performance of machine learning models, allowing model selection and parameter tuning. However, the very process of CV requires random partitioning of the data and so our performance estimates are in fact stochastic, with variability that can be substantial for natural language processing tasks. We demonstrate that these unstable estimates cannot be relied upon for effective parameter tuning. The resulting tuned parameters are highly sensitive to how our data is partitioned, meaning that we often select sub-optimal parameter choices and have serious reproducibility issues. Instead, we propose to use the less variable J-K-fold CV, in which J independent K-fold cross validations are used to assess performance. Our main contributions are extending J-K-fold CV from performance estimation to parameter tuning and investigating how to choose J and K. We argue that variability is more important than bias for effective tuning and so advocate lower choices of K than are typically seen in the NLP literature, instead use the saved computation to increase J. To demonstrate the generality of our recommendations we investigate a wide range of case-studies: sentiment classification (both general and target-specific), part-of-speech tagging and document classification.


AI for text analytics and NLP

#artificialintelligence

There has been a significant growth in the volume and variety of data because of the accumulation of unstructured text data. Companies are now relying on technologies like text analytics and Natural Language Processing (NLP) for making sense of such massively collected data. Text analytics and NLP hold the key to unlocking the business value within these huge data sets. NLP is concerned with making natural language accessible to machines, while text analytics refers to the extraction of useful information from text sources. Today, text analytics and NLP are gradually transforming into a field extremely useful for various business applications, such as competitive analysis, and improving the quality of machine intelligence systems.


Facebook is building a big new $750 million data center in Alabama

#artificialintelligence

Facebook has got big plans for a new $750 million data center in Huntsville, Alabama. On Thursday, the social networking giant announced it was building a new 970,000 square foot facility in Huntsville, a city in the northern part of the US state. "As a growing tech hub, Huntsville seemed like a natural fit for Facebook," the company wrote on a new Facebook post dedicated to the planned data center. "It also provides reliable access to renewable energy, strong local infrastructure, a great set of community partners, and very importantly, an outstanding pool of talent." A Facebook spokesperson confirmed to Business Insider that it is investing $750 million in the project.


New Revelations of Facebook Data Sharing With Device Makers Raises Questions - CPO Magazine

#artificialintelligence

After a New York Times reporter published a bombshell story on how Facebook intentionally shared personal information of its users with over 60 different device makers – including the two largest, Apple and Samsung – in order to create "Facebook-like experiences" on those devices, U.S lawmakers and regulators are asking Facebook for answers. What they see is a clear pattern of abuse, in which Facebook pledges over and over again that it has figured out its data sharing problem, but new revelations continually come to light. In its defense, Facebook says that the data sharing partnerships with mobile device manufacturers were only intended to make it easier for these global device makers to build software incorporating Facebook functionality into their smartphone devices. According to Ime Archibong, Facebook Vice President of Product Partnerships, the data-sharing program was controlled and monitored by Facebook from the very beginning, and all integrations with the device manufacturers were personally approved by Facebook. Moreover, while personal user data might have been shared with these device makers, Facebook says it was only used in order to give users access to Facebook messages and notifications on their mobile phones. But the New York Times report suggests that the data-sharing program went far beyond this.


A Beginner's Guide on Sentiment Analysis with RNN – Towards Data Science

@machinelearnbot

In order to feed this data into our RNN, all input documents must have the same length. We will limit the maximum review length to max_words by truncating longer reviews and padding shorter reviews with a null value (0). We can accomplish this using the pad_sequences() function in Keras. For now, set max_words to 500. We start building our model architecture in the code cell below.


Following Facebook data-sharing revelation, U.S. senator quizzes Alphabet, Twitter on Huawei relationship

The Japan Times

WASHINGTON – A U.S. senator on Thursday is seeking responses from Google parent Alphabet Inc. and Twitter Inc. on whether the U.S. companies have any data-sharing agreements with Chinese vendors, following a disclosure from Facebook Inc. this week. Sen. Mark Warner, a Democrat who is vice chairman of the Intelligence Committee, said in a statement he has written letters to the companies for information on data-sharing agreements, noting that since 2012 "the relationship between the Chinese Communist Party and equipment makers like Huawei and ZTE has been an area of national security concern." Alphabet has said previously it has strategic partnerships with Chinese mobile device manufacturers, including Huawei Technologies Co. Ltd., and Xiaomi, as well as with Chinese technology platform Tencent. It wasn't clear if Twitter has a partnership with any Chinese vendors. Alphabet and Twitter did not immediately respond to questions for comment.


Senator probes Alphabet and Twitter on data-sharing with Chinese firms

Engadget

The New York Times recently revealed that Facebook entered into agreements with at least 60 mobile device companies, giving them access to Facebook user data so that the companies could recreate Facebook-like features. Among those companies are four Chinese firms -- Huawei, Lenovo, Oppo and TCL -- which has spurred some concern among US lawmakers. Today, Senator Mark Warner (D-VA) sent letters to both Alphabet and Twitter, inquiring as to whether they entered into similar data-sharing agreements with any mobile device companies based in China. "Since at least October 2012, when the House Permanent Select Committee on Intelligence released its widely-publicized report, the relationship between the Chinese Communist Party and equipment makers like Huawei and ZTE has been an area of national security concern," wrote Warner. He then goes on to ask both companies if they've had agreements in place similar to Facebook's and if so, whether any Chinese firms like Huawei, ZTE, Lenovo or Xiaomi were included.


Cambridge Analytica ex-boss admits getting Facebook data from researcher

The Japan Times

LONDON – The former head of Cambridge Analytica admitted on Wednesday his firm had received data from the researcher at the center of a scandal over Facebook users' personal details, contradicting previous testimony to lawmakers. Cambridge Analytica, which was hired by Donald Trump in 2016, has denied its work on the U.S. president's successful election campaign made use of data allegedly improperly harvested from around 87 million Facebook users. Former chief Alexander Nix, in earlier testimony to Parliament's media committee, also denied the political consultancy had ever been given data by Aleksandr Kogan, the researcher linked to the scandal. On Wednesday he said it had received data from Kogan. "Of course, the answer to this question should have been'yes,' " Nix said, adding that he thought he was being asked if Cambridge Analytica still held data from the researcher.