Information Extraction
The FBI, SEC and Justice Department Now Want to Know What Facebook Knew About Cambridge Analytica
A federal probe into Facebook's sharing of user data with Cambridge Analytica now involves the FBI, the Securities and Exchange Commission and the Justice Department, the Washington Post reported. Representatives from these agencies have joined the Federal Trade Commission in the inquiry, the newspaper reported, citing five unnamed people familiar with the matter. Those people spoke on condition of anonymity because the probes are not complete. The probe reportedly centers on what Facebook knew in 2015, when it learned that the political data-mining firm Cambridge Analytica had improperly accessed the personal data of tens of millions of Facebook users. Facebook didn't disclose the incident with the political firm, which later worked for the Trump campaign and other Republican candidates, until this March.
Probe Into Facebook's Data Breach Broadens: Washington Post
The emphasis has been on what Facebook has reported publicly about its sharing of information with Cambridge Analytica, whether those representations square with the underlying facts and whether Facebook made sufficiently complete and timely disclosures to the public and investors about the matter, the Washington Post report said.
Learning Semantic Sentence Embeddings using Pair-wise Discriminator
Patro, Badri N., Kurmi, Vinod K., Kumar, Sandeep, Namboodiri, Vinay P.
In this paper, we propose a method for obtaining sentence-level embeddings. While the problem of securing word-level embeddings is very well studied, we propose a novel method for obtaining sentence-level embeddings. This is obtained by a simple method in the context of solving the paraphrase generation task. If we use a sequential encoder-decoder model for generating paraphrase, we would like the generated paraphrase to be semantically close to the original sentence. One way to ensure this is by adding constraints for true paraphrase embeddings to be close and unrelated paraphrase candidate sentence embeddings to be far. This is ensured by using a sequential pair-wise discriminator that shares weights with the encoder that is trained with a suitable loss function. Our loss function penalizes paraphrase sentence embedding distances from being too large. This loss is used in combination with a sequential encoder-decoder network. We also validated our method by evaluating the obtained embeddings for a sentiment analysis task. The proposed method results in semantic embeddings and outperforms the state-of-the-art on the paraphrase generation and sentiment analysis task on standard datasets. These results are also shown to be statistically significant.
NSA Spy Buildings, Facebook Data, and More Security News This Week
It has been, to be quite honest, a fairly bad week, as far as weeks go. But despite the sustained downbeat news, a few good things managed to happen as well. California has passed the strongest digital privacy law in the United States, for starters, which as of 2020 will give customers the right to know what data companies use, and to disallow those companies from selling it. It's just the latest in a string of uncommonly good bits of privacy news, which included last week's landmark Supreme Court decision in Carpenter v. US. That ruling will require law enforcement to get a warrant before accessing cell tower location data.
r/MachineLearning - [D] Don't common sentiment analysis strategies seem unsatisfying?
There's lots of great projects in Reddit in sentiment analysis, but almost all of the work I've seen focuses on individual posts, as if tweets or reddit comments was simply a list of thumbs up and thumbs down about issues. For example, context, which doesn't seem to get much discussion. One very basic example where this is important: a Reddit comment that itself is booing a negative comment is considered negative. Of course, the nested "negative" comment should actually be counted in favor of the original topic. The relevant fields in NLP would be coreference, and possibly other subfields involving semantics.
A Quiz App Exposed 120 Million People's Facebook Data--and Cambridge Analytica Had Nothing to Do With It
Future Tense is a partnership of Slate, New America, and Arizona State University that examines emerging technologies, public policy, and society. The latest chapter in Facebook's data woes involves a quiz app that, until as recently as June, exposed the information of 120 million people who just wanted to know whether they were Cinderella or Elsa. According to De Ceukelaire, beginning as early as the end of 2016, NameTests collected Facebook users' data when they opted to take a quiz, such as "Which Disney Princess Are You?" The app then displayed that data--including names, birthdays, photos, and friends lists--in Javascript files easily accessible by third-party websites. De Ceukelaire writes, "Depending on what quizzes you took, the javascript could leak your Facebook ID, first name, last name, language, gender, date of birth, profile picture, cover photo, currency, devices you use, when your information was lasted updated, your posts and status, your photos and your friends." De Ceukelaire says he "would be surprised if nobody else found this earlier," since the flaw was "really easy to spot," but NameTests said it found no evidence of abuse.
Box expands Skills beta with more AI and machine learning technologies - SiliconANGLE
Box Inc. today expanded a beta test program for its Skills software framework that uses machine learning to make video, audio, image and other files more useful on its content management service. Introduced last year, Box Skills opened up the ability to perform tasks on content such as computer vision for image analysis, video indexing and sentiment analysis from audio using machine learning technologies from IBM Corp.'s Watson, Microsoft Corp's Azure cloud and Google Cloud. Now, Box is expanding beyond the approximately 100 customers in the beta program, adding several new customers a week starting in July on top of the likes of Virgin Trains, Ancestry.com, "As we see more and more content put into Box, there's this opportunity with AI and machine learning to have more intelligence put in around the content," Box Chief Product Officer Jeetu Patel said in an interview. For instance, among the approximately 600 use cases is a large insurance company building a custom skill to label household objects in images and videos automatically in the homeowner insurance policy process.
Wavethrough Vulnerability In Microsoft Edge Could Allow Data Scraping
We all know Microsoft has recently launched a massive'bug fix bundle' where it released patches for around 50 vulnerabilities including the patch for Cortana's Lock Screen Bypass Vulnerability. However, not many know about'all' of these vulnerabilities for which Microsoft released fixes. It was also strange that it released patches together for 50 different bugs. Seems like the team has been silently working out how to solve various issues reported to them over the past months. Now, an independent security researcher has unveiled one such issue.
The word is out. SAS leads in AI.
SAS Visual Text Analytics uses intelligent algorithms and natural language processing (NLP) techniques to automatically extract relationships and patterns within unstructured data, therefore eliminating the need for manual analysis. The NLP tools help users in sentiment analysis, speech to text, natural language understanding and natural language generation. The Forrester report states: "SAS's brand speaks for itself as a leader in advanced analytics; as a result, SAS Visual Text Analytics comes with a number of machine learning models. Users can also leverage other capabilities of the platform, such as forecasting and optimization, to deliver predictive, prescriptive, and actionable analytics."
Harvester of Facebook Data Wants Tighter Controls Over Privacy
Sen. Jerry Moran (R., Kan.), chairman of the Senate's consumer protection subcommittee, said he was considering joining in an effort by Sen. Richard Blumenthal (D., Conn.) to pass a privacy bill of rights in Congress. His comments showed that the risks for big internet companies haven't dissipated since Facebook's scandal involving Cambridge Analytica, a political data consultancy that worked with President Donald Trump's 2016 campaign and obtained data of millions of Facebook users from an app developer, Aleksandr Kogan. Sen. John Thune (R., S.D.), the chairman of the powerful Commerce Committee, added that Facebook "remains under the microscope" and said lawmakers continue to examine potential measures to protect user privacy. But key lawmakers appeared to be far from a consensus on how to proceed. At Tuesday's hearing, Mr. Kogan, a social psychologist and University of Cambridge lecturer, in prepared testimony, called for strengthening the system of obtaining users' consent for subsequent use of their information.