Goto

Collaborating Authors

 Government


Gumbel Counterfactual Generation From Language Models

arXiv.org Artificial Intelligence

Understanding and manipulating the causal generation mechanisms in language models is essential for controlling their behavior. Previous work has primarily relied on techniques such as representation surgery -- e.g., model ablations or manipulation of linear subspaces tied to specific concepts -- to \emph{intervene} on these models. To understand the impact of interventions precisely, it is useful to examine counterfactuals -- e.g., how a given sentence would have appeared had it been generated by the model following a specific intervention. We highlight that counterfactual reasoning is conceptually distinct from interventions, as articulated in Pearl's causal hierarchy. Based on this observation, we propose a framework for generating true string counterfactuals by reformulating language models as a structural equation model using the Gumbel-max trick, which we called Gumbel counterfactual generation. This reformulation allows us to model the joint distribution over original strings and their counterfactuals resulting from the same instantiation of the sampling noise. We develop an algorithm based on hindsight Gumbel sampling that allows us to infer the latent noise variables and generate counterfactuals of observed strings. Our experiments demonstrate that the approach produces meaningful counterfactuals while at the same time showing that commonly used intervention techniques have considerable undesired side effects.


Reasoner Outperforms: Generative Stance Detection with Rationalization for Social Media

arXiv.org Artificial Intelligence

Stance detection is crucial for fostering a human-centric Web by analyzing user-generated content to identify biases and harmful narratives that undermine trust. With the development of Large Language Models (LLMs), existing approaches treat stance detection as a classification problem, providing robust methodologies for modeling complex group interactions and advancing capabilities in natural language tasks. However, these methods often lack interpretability, limiting their ability to offer transparent and understandable justifications for predictions. This study adopts a generative approach, where stance predictions include explicit, interpretable rationales, and integrates them into smaller language models through single-task and multitask learning. We find that incorporating reasoning into stance detection enables the smaller model (FlanT5) to outperform GPT-3.5's zero-shot performance, achieving an improvement of up to 9.57%. Moreover, our results show that reasoning capabilities enhance multitask learning performance but may reduce effectiveness in single-task settings. Crucially, we demonstrate that faithful rationales improve rationale distillation into SLMs, advancing efforts to build interpretable, trustworthy systems for addressing discrimination, fostering trust, and promoting equitable engagement on social media.


VibeCheck: Discover and Quantify Qualitative Differences in Large Language Models

arXiv.org Artificial Intelligence

Large language models (LLMs) often exhibit subtle yet distinctive characteristics in their outputs that users intuitively recognize, but struggle to quantify. These "vibes" -- such as tone, formatting, or writing style -- influence user preferences, yet traditional evaluations focus primarily on the singular axis of correctness. We introduce VibeCheck, a system for automatically comparing a pair of LLMs by discovering identifying traits of a model (vibes) that are well-defined, differentiating, and user-aligned. VibeCheck iteratively discovers vibes from model outputs and then utilizes a panel of LLM judges to quantitatively measure the utility of each vibe. We validate that the vibes generated by VibeCheck align with those found in human discovery and run VibeCheck on pairwise preference data from real-world user conversations with Llama-3-70b vs GPT-4. VibeCheck reveals that Llama has a friendly, funny, and somewhat controversial vibe. These vibes predict model identity with 80% accuracy and human preference with 61% accuracy. Lastly, we run VibeCheck on a variety of models and tasks including summarization, math, and captioning to provide insight into differences in model behavior. VibeCheck discovers vibes like Command X prefers to add concrete intros and conclusions when summarizing in comparison to TNGL, Llama-405b often overexplains its thought process on math problems compared to GPT-4o, and GPT-4 prefers to focus on the mood and emotions of the scene when captioning compared to Gemini-1.5-Flash. Code can be found at https://github.com/lisadunlap/VibeCheck


One world, one opinion? The superstar effect in LLM responses

arXiv.org Artificial Intelligence

As large language models (LLMs) are shaping the way information is shared and accessed online, their opinions have the potential to influence a wide audience. This study examines who the LLMs view as the most prominent figures across various fields, using prompts in ten different languages to explore the influence of linguistic diversity. Our findings reveal low diversity in responses, with a small number of figures dominating recognition across languages (also known as the "superstar effect"). These results highlight the risk of narrowing global knowledge representation when LLMs retrieve subjective information.


AI and the Future of Digital Public Squares

arXiv.org Artificial Intelligence

Two substantial technological advances have reshaped the public square in recent decades: first with the advent of the internet and second with the recent introduction of large language models (LLMs). LLMs offer opportunities for a paradigm shift towards more decentralized, participatory online spaces that can be used to facilitate deliberative dialogues at scale, but also create risks of exacerbating societal schisms. Here, we explore four applications of LLMs to improve digital public squares: collective dialogue systems, bridging systems, community moderation, and proof-of-humanity systems. Building on the input from over 70 civil society experts and technologists, we argue that LLMs both afford promising opportunities to shift the paradigm for conversations at scale and pose distinct risks for digital public squares. We lay out an agenda for future research and investments in AI that will strengthen digital public squares and safeguard against potential misuses of AI.


Exploring Text Representations for Online Misinformation

arXiv.org Artificial Intelligence

Mis- and disinformation, commonly collectively called fake news, continue to menace society. Perhaps, the impact of this age-old problem is presently most plain in politics and healthcare. However, fake news is affecting an increasing number of domains. It takes many different forms and continues to shapeshift as technology advances. Though it arguably most widely spreads in textual form, e.g., through social media posts and blog articles. Thus, it is imperative to thwart the spread of textual misinformation, which necessitates its initial detection. This thesis contributes to the creation of representations that are useful for detecting misinformation. Firstly, it develops a novel method for extracting textual features from news articles for misinformation detection. These features harness the disparity between the thematic coherence of authentic and false news stories. In other words, the composition of themes discussed in both groups significantly differs as the story progresses. Secondly, it demonstrates the effectiveness of topic features for fake news detection, using classification and clustering. Clustering is particularly useful because it alleviates the need for a labelled dataset, which can be labour-intensive and time-consuming to amass. More generally, it contributes towards a better understanding of misinformation and ways of detecting it using Machine Learning and Natural Language Processing.


Drone mystery deepens with Chinese man's troubling Google history after his arrest for 'flying over US base'

Daily Mail - Science & tech

A Chinese man has been arrested for allegedly flying a drone over Vandenberg Space Force Base, as the FBI investigates mysterious drones in New Jersey. Yinpiao Zhou, 39, a Chinese National now living in Brentwood, California, was charged with failure to register an aircraft not providing transportation and violation of national defense airspace. Zhou was arrested Monday at San Francisco International Airport prior to boarding a China-bound flight and made his initial appearance Tuesday in United States District Court in San Francisco. He is in federal custody pending prosecutors' appeal of a federal magistrate judge's decision to release him. No plea was taken and his arraignment is expected to be scheduled in US District Court in Los Angeles in the coming weeks.


More than 20 days into phenomenon, Pentagon still has no answers about origins of mysterious NJ drones

FOX News

Retired Gen. Anthony Tata joined'Fox & Friends' to discuss outrage stemming from the mysterious drones over flying over New Jersey and the Pentagon's lack of clarity over the sightings. More than three weeks after dozens of mysterious drones began popping up in the New Jersey night sky, the public has still been offered no clear insight on what the phenomenon could be. Rep. Jeff Van Drew, R-N.J., suggested the swarms of unmanned aerial vehicles could be from an Iranian "mother ship." "There is no Iranian ship off the coast of the United States, and there's no so-called mother ship launching drones towards the United States," said Pentagon spokesperson Sabrina Singh. She added there is "no evidence" to suggest the drones are "the work of a foreign adversary."


Experts reveal what mystery drones over New Jersey REALLY are... and why Americans should be terrified

Daily Mail - Science & tech

Intelligence analysts have revealed why they believe Russia is behind the mysterious drones invading the skies over New Jersey. US Army general Darryl Williams described a situation that mirrors what has unfolded at American/NATO bases across Europe that are known to supply arms to Ukraine. And retired police lieutenant and intelligence analyst Tim McMillan told DailyMail.com Lt McMillan and other experts have noted that the New Jersey sightings circled around Picatinny Arsenal, home of the US Army's CCDC Armaments Center, which is responsible for manufacturing and supplying Ukraine with artillery ammunition. These experts suggest that Russia could be carrying out an intelligence-gathering mission known as'ferreting', meant to intentionally trigger and test their foreign rival's airspace defense procedures and response time.


Teen deepfake pornography victim warns future generation is 'at risk' if AI crime bill fails

FOX News

High school student Elliston Berry discusses the Take It Down Act, a measure that would force social media companies to remove graphic deepfakes, prevent them from being posted and criminalize the act. Senate lawmakers unanimously passed the bipartisan-led Take It Down Act that would force social media companies to speedily remove sexually explicit deepfakes, prevent them from being posted and criminalize the act. For deepfake pornography victims like 15-year-old Elliston Berry, the measure would be long overdue. The Texas high school student is working with lawmakers to get the bill passed to protect victims like herself. She's inspired by her own story from last year, when she discovered deepfake nude images of herself circulating across social media in a sinister cyber scheme that turned her life upside down.