Media
LEIA: Linguistic Embeddings for the Identification of Affect
Aroyehun, Segun Taofeek, Malik, Lukas, Metzler, Hannah, Haimerl, Nikolas, Di Natale, Anna, Garcia, David
The wealth of text data generated by social media has enabled new kinds of analysis of emotions with language models. These models are often trained on small and costly datasets of text annotations produced by readers who guess the emotions expressed by others in social media posts. This affects the quality of emotion identification methods due to training data size limitations and noise in the production of labels used in model development. We present LEIA, a model for emotion identification in text that has been trained on a dataset of more than 6 million posts with self-annotated emotion labels for happiness, affection, sadness, anger, and fear. LEIA is based on a word masking method that enhances the learning of emotion words during model pre-training. LEIA achieves macro-F1 values of approximately 73 on three in-domain test datasets, outperforming other supervised and unsupervised methods in a strong benchmark that shows that LEIA generalizes across posts, users, and time periods. We further perform an out-of-domain evaluation on five different datasets of social media and other sources, showing LEIA's robust performance across media, data collection methods, and annotation schemes. Our results show that LEIA generalizes its classification of anger, happiness, and sadness beyond the domain it was trained on. LEIA can be applied in future research to provide better identification of emotions in text from the perspective of the writer. The models produced for this article are publicly available at https://huggingface.co/LEIA
Trust and Reliance in Consensus-Based Explanations from an Anti-Misinformation Agent
Ueno, Takane, Kim, Yeongdae, Oura, Hiroki, Seaborn, Katie
The illusion of consensus occurs when people believe there is consensus across multiple sources, but the sources are the same and thus there is no "true" consensus. We explore this phenomenon in the context of an AI-based intelligent agent designed to augment metacognition on social media. Misinformation, especially on platforms like Twitter, is a global problem for which there is currently no good solution. As an explainable AI (XAI) system, the agent provides explanations for its decisions on the misinformed nature of social media content. In this late-breaking study, we explored the roles of trust (attitude) and reliance (behaviour) as key elements of XAI user experience (UX) and whether these influenced the illusion of consensus. Findings show no effect of trust, but an effect of reliance on consensus-based explanations. This work may guide the design of anti-misinformation systems that use XAI, especially the user-centred design of explanations.
A Group-Specific Approach to NLP for Hate Speech Detection
Automatic hate speech detection is an important yet complex task, requiring knowledge of common sense, stereotypes of protected groups, and histories of discrimination, each of which may constantly evolve. In this paper, we propose a group-specific approach to NLP for online hate speech detection. The approach consists of creating and infusing historical and linguistic knowledge about a particular protected group into hate speech detection models, analyzing historical data about discrimination against a protected group to better predict spikes in hate speech against that group, and critically evaluating hate speech detection models through lenses of intersectionality and ethics. We demonstrate this approach through a case study on NLP for detection of antisemitic hate speech. The case study synthesizes the current English-language literature on NLP for antisemitism detection, introduces a novel knowledge graph of antisemitic history and language from the 20th century to the present, infuses information from the knowledge graph into a set of tweets over Logistic Regression and uncased DistilBERT baselines, and suggests that incorporating context from the knowledge graph can help models pick up subtle stereotypes.
The Thing About Hostage Taking Is Who Pays The Ransom
This week, David Plotz, John Dickerson, and Emily Bazelon discuss the $787.5 million settlement of the Dominion Voting v. Fox News defamation lawsuit; the political game being played with raising the U.S. debt ceiling; and the Russian detention of American journalist Evan Gershkovich. Here are some notes and references from this week's show: Matthew Iglesias for Slow Boring: "Medicaid work requirements are cruel and pointless" John Dickerson for CBS News Prime Time: "U.S. ambassador says she visited detained Wall Street Journal reporter" Drew Hinshaw, Joe Parkinson, and Brett Forrest for the Wall Street Journal: "'You Are Completely Alone': Inside the Infamous Russian Prison Holding Evan Gershkovich" Centers for Disease Control and Prevention: "What Everyone Should Know about the Shingles Vaccine (Shingrix)" Carrie Blazina and Drew Desilver for the Pew Research Center: "House gets younger, Senate gets older: A look at the age and generation of lawmakers in the 118th Congress" Here are this week's chatters: Emily: Julie Bosman, Mitch Smith, Jesse McKinley, and Jay Root for the New York Times: "Hundreds of Miles Apart, Separate Shootings Follow Wrong Turns" and Timothy Bella for the Washington Post: "Cheerleaders leaving practice were shot after one got in wrong car, teen says" John: Ellie Zolfagharifard for the Daily Mail: "'Here there be robots': Artist draws stunning medieval map of Mars showing off its huge craters and vast canyons"; Mars and its Canals by Percival Lowell; and Kaushik Patowary for Amusing Planet: "How Astronomer Percival Lowell Mistook His Own Eye For Spokes on Venus" David: City Cast DC podcast: "D.C.'s Rat-Hunting Dogs And Other Rat Solutions" (Host Bridget Todd, Producer Julia Karron) Listener chatter from Nancy Hall: Joe Mahr and Megan Crepeau for the Chicago Tribune: "Stalled Justice: Delays in the Cook County courts" For this week's Slate Plus bonus segment, Emily, John, and David discuss the dilemma posed by the months-long absence of Dianne Feinstein from the U.S. Senate. In the next Gabfest Reads, David talks with Washington Post columnist Alexandra Petri about her latest book, Alexandra Petri's US History: Important American Documents (I Made Up).
AI researchers claim Google, '60 Minutes' spread 'disinformation' in recent interview: 'Still bulls---'
Sundar Pichai told '60 Minutes' that the state of the technology is still somewhat of a black box to researchers. Researchers are accusing Google and CBS News of overestimating the capabilities of artificial intelligence (AI) following an interview between the Alphabet CEO Sundar Pichai and "60 Minutes." During the recent interview, Pichai claimed that AI programs developed by Google had displayed "emergent properties," or the ability to learn unexpected skills they were not trained on, puzzling researchers. For example, Google tech executive James Manyika claimed the company's AI had learned the language of Bengali without significant implementation of the information beforehand. "We discovered that with very few amounts of prompting in Bengali," Manyika said, "it can now translate all of Bengali."
Stack Overflow Will Charge AI Giants for Training Data
Developing the AI systems behind tools such as ChatGPT and the image generator Dall-E costs hundreds of millions of dollars--and it's about to get more expensive. OpenAI, Google, and other companies building large-scale AI projects have traditionally paid nothing for much of their training data, scraping it from the web. But Stack Overflow, a popular internet forum for computer programming help, plans to begin charging large AI developers as soon as the middle of this year for access to the 50 million questions and answers on its service, CEO Prashanth Chandrasekar says. The site has more than 20 million registered users. Stack Overflow's decision to seek compensation from companies tapping its data, part of a broader generative AI strategy, has not been previously reported. It follows an announcement by Reddit this week that it will begin charging some AI developers to access its own content starting in June.
What is Snapchat AI? Instant messaging app rolls out its ChatGPT-powered AI chatbot
Fox News correspondent CB Cotton has the latest on calls for accountability for social media apps after parents say Snapchat helped facilitate drug sales on'Special Report.' Snapchat rolled out its ChatGPT-powered AI chatbot called "My AI" on Wednesday. The feature was introduced in February and was previously only available to Snapchat subscribers, but will now be available to all, the latest move by a social media company in the quickly evolving artificial intelligence (AI) race. "Snapchat subscribers have been loving My AI, our AI-powered chatbot, sending nearly 2 million chat messages per day to learn more about movies, sports, pets, and the world around them," Snap "Today, we announced My AI is rolling out to Snapchatters globally, now with brand new features," Snapchat's parent company, Snap, wrote in a press release. "My AI is an experimental chatbot," a message to Snapchat users on Wednesday read.
The Drake AI Song Is Just the Tip of the Iceberg
The notion that Drake allegedly uses a ghostwriter to write his rhymes is a conspiracy that has haunted the rapper for years. He told Genius that he doesn't lean on ghostwriters, saying, "Any song that really, really did damage for me, I wrote every single lyric." The rumors were also the subject of his famous feud with rapper Meek Mill, spawning his pair of diss tracks, "Charged Up" and "Back to Back." But now, a more ominous presence has appeared on social media platforms to actually ghostwrite a Drake song--sort of. Last weekend a TikTok creator by the name of @ghostwriter977 uploaded a video in which they premiered an AI-generated "Drake" track titled "Heart on My Sleeve" with a faux-assist from a similarly AI-generated The Weeknd.
Fresh concerns raised over sources of training material for AI systems
Fresh fears have been raised about the training material used for some of the largest and most powerful artificial intelligence models, after several investigations exposed the fascist, pirated and malicious sources from which the data is harvested. One such dataset is the Colossal Clean Crawled Corpus, or C4, assembled by Google from more than 15m websites and used to train both the search engine's LaMDA AI as well as Meta's GPT competitor, LLaMA. The dataset is public, but its scale has made it difficult to examine the contents: it is supposedly a "clean" version of a more expansive dataset, Common Crawl, with "noisy" content, offensive language and racist slurs removed from the material. But an investigation by the Washington Post reveals that C4's "cleanliness" is only skin deep. While it draws on websites such as the Guardian – which makes up 0.05% of the entire dataset - and Wikipedia, as well as large databases such as Google Patents and the scientific journal hub PLOS, it also contains less reputable sites. The white nationalist site VDARE is in the database, one of the 1,000 largest sites, as is the far-right news site Breitbart.
UK auction: Rare 'Magic: The Gathering' cards will be up for bid, expected to fetch almost $200K
A collection of extremely rare Magic: The Gathering cards are expected to fetch as much as $180,000 when they are auctioned off on Friday, April 21. Included among the cards that are going up for auction is a complete 310-card set of "Legends," the third expansion pack sold by Magic: The Gathering, which was released in June 1994. The auction will also feature a factory-sealed unlimited starter deck, which is expected to sell for at least £10,000-12,000, or about $15,000 U.S. dollars. A starter deck is a random collection of 60 cards from a set. Magic: The Gathering is a collectible card game produced by Wizards of the Coast.