Government
Evidencing Unauthorized Training Data from AI Generated Content using Information Isotopes
Tao, Qi, Jinhua, Yin, Dongqi, Cai, Yueqi, Xie, Huili, Wang, Zhiyang, Hu, Peiru, Yang, Guoshun, Nan, Zhili, Zhou, Shangguang, Wang, Lingjuan, Lyu, Yongfeng, Huang, Nicholas, Lane
In light of scaling laws, many AI institutions are intensifying efforts to construct advanced AIs on extensive collections of high-quality human data. However, in a rush to stay competitive, some institutions may inadvertently or even deliberately include unauthorized data (like privacy- or intellectual property-sensitive content) for AI training, which infringes on the rights of data owners. Compounding this issue, these advanced AI services are typically built on opaque cloud platforms, which restricts access to internal information during AI training and inference, leaving only the generated outputs available for forensics. Thus, despite the introduction of legal frameworks by various countries to safeguard data rights, uncovering evidence of data misuse in modern opaque AI applications remains a significant challenge. In this paper, inspired by the ability of isotopes to trace elements within chemical reactions, we introduce the concept of information isotopes and elucidate their properties in tracing training data within opaque AI systems. Furthermore, we propose an information isotope tracing method designed to identify and provide evidence of unauthorized data usage by detecting the presence of target information isotopes in AI generations. We conduct experiments on ten AI models (including GPT-4o, Claude-3.5, and DeepSeek) and four benchmark datasets in critical domains (medical data, copyrighted books, and news). Results show that our method can distinguish training datasets from non-training datasets with 99\% accuracy and significant evidence (p-value$<0.001$) by examining a data entry equivalent in length to a research paper. The findings show the potential of our work as an inclusive tool for empowering individuals, including those without expertise in AI, to safeguard their data rights in the rapidly evolving era of AI advancements and applications.
Scaling Laws for Emulation of Stellar Spectra
Rรณลผaลski, Tomasz, Ting, Yuan-Sen
Neural network-based emulators for the inference of stellar parameters and elemental abundances represent an increasingly popular methodology in modern spectroscopic surveys. However, these approaches are often constrained by their emulation precision and domain transfer capabilities. Greater generalizability has previously been achieved only with significantly larger model architectures, as demonstrated by Transformer-based models in natural language processing. This observation aligns with neural scaling laws, where model performance predictably improves with increased model size, computational resources allocated to model training, and training data volume. In this study, we demonstrate that these scaling laws also apply to Transformer-based spectral emulators in astronomy. Building upon our previous work with TransformerPayne and incorporating Maximum Update Parametrization techniques from natural language models, we provide training guidelines for scaling models to achieve optimal performance. Our results show that within the explored parameter space, clear scaling relationships emerge. These findings suggest that optimal computational resource allocation requires balanced scaling. Specifically, given a tenfold increase in training compute, achieving an optimal seven-fold reduction in mean squared error necessitates an approximately 2.5-fold increase in dataset size and a 3.8-fold increase in model size. This study establishes a foundation for developing spectral foundational models with enhanced domain transfer capabilities.
Masks and Mimicry: Strategic Obfuscation and Impersonation Attacks on Authorship Verification
Alperin, Kenneth, Leekha, Rohan, Uchendu, Adaku, Nguyen, Trang, Medarametla, Srilakshmi, Capote, Carlos Levya, Aycock, Seth, Dagli, Charlie
The increasing use of Artificial Intelligence (AI) technologies, such as Large Language Models (LLMs) has led to nontrivial improvements in various tasks, including accurate authorship identification of documents. However, while LLMs improve such defense techniques, they also simultaneously provide a vehicle for malicious actors to launch new attack vectors. To combat this security risk, we evaluate the adversarial robustness of authorship models (specifically an authorship verification model) to potent LLM-based attacks. These attacks include untargeted methods - \textit{authorship obfuscation} and targeted methods - \textit{authorship impersonation}. For both attacks, the objective is to mask or mimic the writing style of an author while preserving the original texts' semantics, respectively. Thus, we perturb an accurate authorship verification model, and achieve maximum attack success rates of 92\% and 78\% for both obfuscation and impersonation attacks, respectively.
LLM Benchmarking with LLaMA2: Evaluating Code Development Performance Across Multiple Programming Languages
Diehl, Patrick, Nader, Nojoud, Moraru, Maxim, Brandt, Steven R.
Large Language Models (LLMs) have made significant advances in various code-related tasks, particularly in generating source code from natural language descriptions (Zhao et al. (2023); Chang et al. (2024)). Their effectiveness is primarily driven by their extensive number of model parameters, the use of large and diverse datasets, and the immense computational resources employed during training (Kaplan et al. (2020)). These models are typically trained on vast corpora sourced from the web. LLMs are capable of capturing intricate patterns, linguistic subtleties, and semantic relationships. A wide range of models are available for code generation. There are general-purpose models like ChatGPT (Ouyang et al. (2022)), GPT -4 (Achiam et al. (2023)), and LLaMA (Touvron et al. (2023a)) which are designed for a broad range of applications, as well as specialized models such as StarCoder, Code LLaMA (Roziere et al. (2023)), DeepSeek-Coder, and Code Gemma that are optimized for code-related tasks. The integration of code generation with the latest advances in LLM technology is now an essential tool for many businesses, as well as an essential target for LLM developers as programming languages are considered to be different dialects of natural language (Athiwaratkun et al. (2022)).
Words as Bridges: Exploring Computational Support for Cross-Disciplinary Translation Work
Bao, Calvin, Shiue, Yow-Ting, Carpuat, Marine, Chan, Joel
Scholars often explore literature outside of their home community of study. This exploration process is frequently hampered by field-specific jargon. Past computational work often focuses on supporting translation work by removing jargon through simplification and summarization; here, we explore a different approach that preserves jargon as useful bridges to new conceptual spaces. Specifically, we cast different scholarly domains as different language-using communities, and explore how to adapt techniques from unsupervised cross-lingual alignment of word embeddings to explore conceptual alignments between domain-specific word embedding spaces.We developed a prototype cross-domain search engine that uses aligned domain-specific embeddings to support conceptual exploration, and tested this prototype in two case studies. We discuss qualitative insights into the promises and pitfalls of this approach to translation work, and suggest design insights for future interfaces that provide computational support for cross-domain information seeking.
Agent-based Modeling meets the Capability Approach for Human Development: Simulating Homelessness Policy-making
Aguilera, Alba, Osman, Nardine, Curto, Georgina
The global rise in homelessness calls for urgent and alternative policy solutions. Non-profits and governmental organizations alert about the many challenges faced by people experiencing homelessness (PEH), which include not only the lack of shelter but also the lack of opportunities for personal development. In this context, the capability approach (CA), which underpins the United Nations Sustainable Development Goals (SDGs), provides a comprehensive framework to assess inequity in terms of real opportunities. This paper explores how the CA can be combined with agent-based modelling and reinforcement learning. The goals are: (1) implementing the CA as a Markov Decision Process (MDP), (2) building on such MDP to develop a rich decision-making model that accounts for more complex motivators of behaviour, such as values and needs, and (3) developing an agent-based simulation framework that allows to assess alternative policies aiming to expand or restore people's capabilities. The framework is developed in a real case study of health inequity and homelessness, working in collaboration with stakeholders, non-profits and domain experts. The ultimate goal of the project is to develop a novel agent-based simulation framework, rooted in the CA, which can be replicated in a diversity of social contexts to assess policies in a non-invasive way.
Accenture-NVS1: A Novel View Synthesis Dataset
Sugg, Thomas, O'Brien, Kyle, Poudel, Lekh, Dumouchelle, Alex, Jou, Michelle, Bosch, Marc, Ramanan, Deva, Narasimhan, Srinivasa, Tulsiani, Shubham
This paper introduces ACC-NVS1, a specialized dataset designed for research on Novel View Synthesis specifically for airborne and ground imagery. Data for ACC-NVS1 was collected in Austin, TX and Pittsburgh, PA in 2023 and 2024. The collection encompasses six diverse real-world scenes captured from both airborne and ground cameras, resulting in a total of 148,000 images. ACC-NVS1 addresses challenges such as varying altitudes and transient objects. This dataset is intended to supplement existing datasets, providing additional resources for comprehensive research, rather than serving as a benchmark.
Manipulation and the AI Act: Large Language Model Chatbots and the Danger of Mirrors
Large Language Model chatbots are increasingly taking the form and visage of human beings, adapting human faces, names, voices, personalities, and quirks, including those of celebrities and well-known political figures. Personifying AI chatbots could foreseeably increase their trust with users. However, it could also make them more capable of manipulation, by creating the illusion of a close and intimate relationship with an artificial entity. The European Commission has finalized the AI Act, with the EU Parliament making amendments banning manipulative and deceptive AI systems that cause significant harm to users. Although the AI Act covers harms that accumulate over time, it is unlikely to prevent harms associated with prolonged discussions with AI chatbots. Specifically, a chatbot could reinforce a person's negative emotional state over weeks, months, or years through negative feedback loops, prolonged conversations, or harmful recommendations, contributing to a user's deteriorating mental health.
Optimizing Influence Campaigns: Nudging under Bounded Confidence
Influence campaigns in online social networks are often run by organizations, political parties, and nation states to influence large audiences. These campaigns are employed through the use of agents in the network that share persuasive content. Yet, their impact might be minimal if the audiences remain unswayed, often due to the bounded confidence phenomenon, where only a narrow spectrum of viewpoints can influence them. Here we show that to persuade under bounded confidence, an agent must nudge its targets to gradually shift their opinions. Using a control theory approach, we show how to construct an agent's nudging policy under the bounded confidence opinion dynamics model and also how to select targets for multiple agents in an influence campaign on a social network. Simulations on real Twitter networks show that a multi-agent nudging policy can shift the mean opinion, decrease opinion polarization, or even increase it. We find that our nudging based policies outperform other common techniques that do not consider the bounded confidence effect. Finally, we show how to craft prompts for large language models, such as ChatGPT, to generate text-based content for real nudging policies. This illustrates the practical feasibility of our approach, allowing one to go from mathematical nudging policies to real social media content.
A deal in the desert? US and Ukraine meet ahead of Russia ceasefire talks
"I feel that he (Putin) wants peace," said President Trump's personal envoy Steve Witkoff, adding: "I think that you're going to see in Saudi Arabia on Monday some real progress." Yet Dmitry Peskov, the Kremlin spokesman has dampened expectations. "We are only at the beginning of this path," he told Russian state TV. Kyiv suffered one of its heaviest attacks from Russian drones on Saturday night, with three people killed, including a five-year-old girl. "We need to push Putin to give a real order to stop the strikes," said Ukraine's President Volodymyr Zelensky in his evening address on Sunday.