Goto

Collaborating Authors

 Information Extraction


Intimate details of 3 MILLION users were exposed in a new Facebook data leak

Daily Mail - Science & tech

Three million Facebook users had their most intimate details exposed as a new data protection scandal hits the social media platform. In the latest of a string of security breaches, a report from New Scientist has revealed a popular personality app insufficiently protected the'anonymous' data of participants. The quiz, called myPersonality, collected highly sensitive data, including psychometric test results that revealed how neurotic or extrovert an individual was. The investigation found the information was poorly protected for four years and gaining access to was relatively easy. Three million of Mark Zuckerberg's Facebook users had intimate details exposed as a new data protection scandal has hit the social media platform.


Huge new Facebook data leak exposed intimate details of 3m users

New Scientist

Data from millions of Facebook users who used a popular personality app, including their answers to intimate questionnaires, was left exposed online for anyone to access, a New Scientist investigation has found. Academics at the University of Cambridge distributed the data from the personality quiz app myPersonality to hundreds of researchers via a website with insufficient security provisions, which led to it being left vulnerable to access for four years. Gaining access illicitly was relatively easy. The data was highly sensitive, revealing personal details of Facebook users, such as the results of psychological tests. It was meant to be stored and shared anonymously, however such poor precautions were taken that deanonymising would not be hard.


Facebook Suspends 200 Apps Amidst Data Privacy Investigation. And More Could Be Coming

TIME - Tech

At least 200 apps have been suspended from Facebook amidst a data privacy investigation launched by Mark Zuckerberg after the Cambridge Analytica scandal in March. On Monday, Facebook announced its internal investigation was in "full swing" -- with teams delving into thousands of apps that are connected to Facebook, according to a statement released by Ime Archibong, vice president of Facebook's product partnerships. Facebook's investigation has already led to the suspension of around 200 apps which will be analyzed to see "whether they did in fact misuse any data." Archibong said the second phase of the investigation involves looking into whether there is evidence that the suspended apps or other apps misused data. If an app misled users in how their data was being used, it could be banned from Facebook.


MojiTalk: Generating Emotional Responses at Scale

arXiv.org Artificial Intelligence

Generating emotional language is a key step towards building empathetic natural language processing agents. However, a major challenge for this line of research is the lack of large-scale labeled training data, and previous studies are limited to only small sets of human annotated sentiment labels. Additionally, explicitly controlling the emotion and sentiment of generated text is also difficult. In this paper, we take a more radical approach: we exploit the idea of leveraging Twitter data that are naturally labeled with emojis. More specifically, we collect a large corpus of Twitter conversations that include emojis in the response, and assume the emojis convey the underlying emotions of the sentence. We then introduce a reinforced conditional variational encoder approach to train a deep generative model on these conversations, which allows us to use emojis to control the emotion of the generated text. Experimentally, we show in our quantitative and qualitative analyses that the proposed models can successfully generate high-quality abstractive conversation responses in accordance with designated emotions.


Python Algo Trading: Market Neutral Hedge Fund Strategy

@machinelearnbot

Update 23 Aug 2017: Do note that Quantopian platform will no longer support third party broker integration. Please see their website under forum. The title of the post is "Phasing Out Brokerage Integrations". This course provides you with the tools that top hedge funds used. These institutional tools include but are not limited to market data, fundamental data, sentiment analysis data, and more.


Real-World Data Mining: Applied Business Analytics and Decision Making

@machinelearnbot

Use the latest data mining best practices to enable timely, actionable, evidence-based decision making throughout your organization! Real-World Data Mining demystifies current best practices, showing how to use data mining to uncover hidden patterns and correlations, and leverage these to improve all aspects of business performance. Drawing on extensive experience as a researcher, practitioner, and instructor, Dr. Dursun Delen delivers an optimal balance of concepts, techniques and applications. Without compromising either simplicity or clarity, he provides enough technical depth to help readers truly understand how data mining technologies work. Coverage includes: processes, methods, techniques, tools, and metrics; the role and management of data; text and web mining; sentiment analysis; and Big Data integration.


Learning Domain-Sensitive and Sentiment-Aware Word Embeddings

arXiv.org Artificial Intelligence

Word embeddings have been widely used in sentiment classification because of their efficacy for semantic representations of words. Given reviews from different domains, some existing methods for word embeddings exploit sentiment information, but they cannot produce domain-sensitive embeddings. On the other hand, some other existing methods can generate domain-sensitive word embeddings, but they cannot distinguish words with similar contexts but opposite sentiment polarity. We propose a new method for learning domain-sensitive and sentiment-aware embeddings that simultaneously capture the information of sentiment semantics and domain sensitivity of individual words. Our method can automatically determine and produce domain-common embeddings and domain-specific embeddings. The differentiation of domain-common and domain-specific words enables the advantage of data augmentation of common semantics from multiple domains and capture the varied semantics of specific words from different domains at the same time. Experimental results show that our model provides an effective way to learn domain-sensitive and sentiment-aware word embeddings which benefit sentiment classification at both sentence level and lexicon term level.


Cambridge Analytica's Facebook data models survived until 2017

Engadget

Facebook may have succeeded in getting Cambridge Analytica to delete millions of users' data in January 2016, but the information based on that data appears to have survived for much longer. The Guardian has obtained leaked emails suggesting that Cambridge Analytica avoided explicitly agreeing to delete the derivatives of that data, such as predictive personality models. Former employees claimed the company kept that data modelling in a "hidden corner" of a server until an audit in March 2017 (prompted by an Observer journalist's investigation), and it only certified that it had scrubbed the data models in April 2017 -- half a year after the US presidential election. In a response to the Guardian, a Cambridge Analytica spokesperson denied that there was a "secret cache," and said that it had started looking for and deleting derivatives of that data after the initial wipe, finishing in April 2017. It was a "lengthy process," the company claimed.


Cambridge Analytica kept Facebook data models through US election

#artificialintelligence

Facebook's failure to compel Cambridge Analytica to delete all traces of data from its servers – including any "derivatives" – enabled the company to retain predictive models derived from millions of social media profiles throughout the US presidential election, the Guardian can reveal. Leaked emails reveal that when Cambridge Analytica told Facebook almost a year before the election that it had deleted data harvested from tens of millions of Facebook users, it stopped short of agreeing to also erase derivatives of the data. The correspondence, obtained by the Guardian, also raises questions about the accuracy of the testimony that Facebook's chief executive, Mark Zuckerberg, gave to the US Congress last month. Derivatives of data, which can include predictive models, or clusters of populations in psychological groupings, can be highly valuable to companies involved in micro-targeting advertisements to voters. Data scientists say such models and analysis are often more valuable than underlying raw data.


Qualitative Data Science: Using RQDA to analyse interviews

#artificialintelligence

Qualitative data science sounds like a contradiction in terms. Data scientists generally solve problems using numerical solutions. Even the analysis of text is reduced to a numerical problem using Markov chains, topic analysis, sentiment analysis and other mathematical tools. Scientists and professionals consider numerical methods the gold standard of analysis. There is, however, a price to pay when relying on numbers alone.