Media
High-level Approaches to Detect Malicious Political Activity on Twitter
Our work represents another step into the detection and prevention of these ever-more present political manipulation efforts. We, therefore, start by focusing on understanding what the state-of-the-art approaches lack -- since the problem remains, this is a fair assumption. We find concerning issues within the current literature and follow a diverging path. Notably, by placing emphasis on using data features that are less susceptible to malicious manipulation and also on looking for high-level approaches that avoid a granularity level that is biased towards easy-to-spot and low impact cases. We designed and implemented a framework -- Twitter Watch -- that performs structured Twitter data collection, applying it to the Portuguese Twittersphere. We investigate a data snapshot taken on May 2020, with around 5 million accounts and over 120 million tweets (this value has since increased to over 175 million). The analyzed time period stretches from August 2019 to May 2020, with a focus on the Portuguese elections of October 6th, 2019. However, the Covid-19 pandemic showed itself in our data, and we also delve into how it affected typical Twitter behavior. We performed three main approaches: content-oriented, metadata-oriented, and network interaction-oriented. We learn that Twitter's suspension patterns are not adequate to the type of political trolling found in the Portuguese Twittersphere -- identified by this work and by an independent peer - nor to fake news posting accounts. We also surmised that the different types of malicious accounts we independently gathered are very similar both in terms of content and interaction, through two distinct analysis, and are simultaneously very distinct from regular accounts.
Chord Embeddings: Analyzing What They Capture and Their Role for Next Chord Prediction and Artist Attribute Prediction
Lahnala, Allison, Kambhatla, Gauri, Peng, Jiajun, Whitehead, Matthew, Minnehan, Gillian, Guldan, Eric, Kummerfeld, Jonathan K., รamcฤฑ, Anฤฑl, Mihalcea, Rada
Natural language processing methods have been applied in a variety of music studies, drawing the connection between music and language. In this paper, we expand those approaches by investigating \textit{chord embeddings}, which we apply in two case studies to address two key questions: (1) what musical information do chord embeddings capture?; and (2) how might musical applications benefit from them? In our analysis, we show that they capture similarities between chords that adhere to important relationships described in music theory. In the first case study, we demonstrate that using chord embeddings in a next chord prediction task yields predictions that more closely match those by experienced musicians. In the second case study, we show the potential benefits of using the representations in tasks related to musical stylometrics.
Controlling Hallucinations at Word Level in Data-to-Text Generation
Rebuffel, Clรฉment, Roberti, Marco, Soulier, Laure, Scoutheeten, Geoffrey, Cancelliere, Rossella, Gallinari, Patrick
Data-to-Text Generation (DTG) is a subfield of Natural Language Generation aiming at transcribing structured data in natural language descriptions. The field has been recently boosted by the use of neural-based generators which exhibit on one side great syntactic skills without the need of hand-crafted pipelines; on the other side, the quality of the generated text reflects the quality of the training data, which in realistic settings only offer imperfectly aligned structure-text pairs. Consequently, state-of-art neural models include misleading statements - usually called hallucinations - in their outputs. The control of this phenomenon is today a major challenge for DTG, and is the problem addressed in the paper. Previous work deal with this issue at the instance level: using an alignment score for each table-reference pair. In contrast, we propose a finer-grained approach, arguing that hallucinations should rather be treated at the word level. Specifically, we propose a Multi-Branch Decoder which is able to leverage word-level labels to learn the relevant parts of each training instance. These labels are obtained following a simple and efficient scoring procedure based on co-occurrence analysis and dependency parsing. Extensive evaluations, via automated metrics and human judgment on the standard WikiBio benchmark, show the accuracy of our alignment labels and the effectiveness of the proposed Multi-Branch Decoder. Our model is able to reduce and control hallucinations, while keeping fluency and coherence in generated texts. Further experiments on a degraded version of ToTTo show that our model could be successfully used on very noisy settings.
Hierarchical Multi-head Attentive Network for Evidence-aware Fake News Detection
To detect fake news, researchers proposed to use The proliferation of biased news, misleading linguistics and textual content (Castillo et al., 2011; claims, disinformation and fake news has caused Zhao et al., 2015; Liu et al., 2015). Since textual heightened negative effects on modern society in claims are usually deliberately written to deceive various domains ranging from politics, economics readers, it is hard to detect fake news by solely to public health. A recent study showed that maliciously relying on the content claims. Therefore, multiple fabricated and partisan stories possibly works utilized other signals such as temporal caused citizens' misperception about political candidates spreading patterns (Liu and Wu, 2018), network (Allcott and Gentzkow, 2017) during the structures (Wu and Liu, 2018; Vo and Lee, 2018; 2016 U.S. presidential elections. In economics, the Shu et al., 2020) and users' feedbacks (Vo and spread of fake news has manipulated stock price Lee, 2019; Shu et al., 2019; Vo and Lee, 2020a).
Inside the mind of Jeff Bezos
The first thing I ever bought on Amazon was an edutainment DVD for babies. I don't recall making the purchase, but the data is unequivocal on this point: on 14 November 2004, I bought Baby Einstein: Baby Noah โ Animal Expedition for the sum of ยฃ7.85. My nearest guess is that I got it as a Christmas present for my nephew, who would at that point have been one year old, and at the very peak of his interest in finger-puppet animals who cavort to xylophone arrangements of Beethoven. This was swiftly followed by three more DVD purchases I have no memory of making. Strangely, I bought nothing at all from Amazon the following year, and then, in 2006, I embarked on a PhD and started ramping up my acquisition of the sort of books that were not easily to be found in brick-and-mortar establishments. Everything ever published by the American novelist Nicholson Baker. I know these things because I recently spent a desultory morning clicking through all 16 years of my Amazon purchase history. Seeing all those hundreds of items bought and delivered, many of them long since forgotten, was a vaguely melancholy experience. I experienced an estranged recognition, as if reading an avant-garde biography of myself, ghost-written by an algorithm. From the bare facts of the things I once bought, I began to reconstruct where I was in life, and what I was doing at the time, and what I was (or wanted to be) interested in. And yet an essential mystery endured.
Microsoft opens limited access to its neural text-to-speech AI
Microsoft is opening up limited access to a text-to-speech AI called Custom Neural Voice, which allows developers to create custom synthetic voices. The tech is part of an Azure AI service called Speech. Companies can use the tech for things like voice-powered smart assistants and devices, chatbots, online learning and reading audiobooks or news. They'll have to apply for access and gain approval from Microsoft before they can harness Custom Neural Voice. The tech can deliver more natural-sounding voices than many other text-to-speech services, according to Microsoft.
Problematic Machine Behavior: A Systematic Literature Review of Algorithm Audits
While algorithm audits are growing rapidly in commonality and public importance, relatively little scholarly work has gone toward synthesizing prior work and strategizing future research in the area. This systematic literature review aims to do just that, following PRISMA guidelines in a review of over 500 English articles that yielded 62 algorithm audit studies. The studies are synthesized and organized primarily by behavior (discrimination, distortion, exploitation, and misjudgement), with codes also provided for domain (e.g. search, vision, advertising, etc.), organization (e.g. Google, Facebook, Amazon, etc.), and audit method (e.g. sock puppet, direct scrape, crowdsourcing, etc.). The review shows how previous audit studies have exposed public-facing algorithms exhibiting problematic behavior, such as search algorithms culpable of distortion and advertising algorithms culpable of discrimination. Based on the studies reviewed, it also suggests some behaviors (e.g. discrimination on the basis of intersectional identities), domains (e.g. advertising algorithms), methods (e.g. code auditing), and organizations (e.g. Twitter, TikTok, LinkedIn) that call for future audit attention. The paper concludes by offering the common ingredients of successful audits, and discussing algorithm auditing in the context of broader research working toward algorithmic justice.
Self-Supervised Claim Identification for Automated Fact Checking
Pathak, Archita, Shaikh, Mohammad Abuzar, Srihari, Rohini
We propose a novel, attention-based self-supervised approach to identify "claim-worthy" sentences in a fake news article, an important first step in automated fact-checking. We leverage "aboutness" of headline and content using attention mechanism for this task. The identified claims can be used for downstream task of claim verification for which we are releasing a benchmark dataset of manually selected compelling articles with veracity labels and associated evidence. This work goes beyond stylistic analysis to identifying content that influences reader belief. Experiments with three datasets show the strength of our model. Data and code available at https://github.com/architapathak/Self-Supervised-ClaimIdentification
'Users' is a fascinating meditation on life and parenting in the digital age
One of the earliest images in Natalia Almada's virtuoso documentary Users is of an infant, tightly wrapped and strapped to a Snoo smart crib, robotically being rocked to sleep to the sound of manufactured white noise. By recreating many of the sensations of being in the womb, the Snoo has become a popular gadget for new parents who need help tucking their little ones in. In many ways, it's the pinnacle of a smart gadget: Developed by Dr. Harvey Karp, with product design by the renowned Yves Behar, the Snoo solves a problem that parents have faced for millennia. But what do we lose if a robot can automatically soothe a crying baby, effectively replacing a nurturing parent. That's the question at the heart of Users, which premiered at the Sundance Film Festival this week.