Goto

Collaborating Authors

 Generative AI


Fundamental Limitations in Defending LLM Finetuning APIs

arXiv.org Artificial Intelligence

LLM developers have imposed technical interventions to prevent fine-tuning misuse attacks, attacks where adversaries evade safeguards by fine-tuning the model using a public API. Previous work has established several successful attacks against specific fine-tuning API defences. In this work, we show that defences of fine-tuning APIs that seek to detect individual harmful training or inference samples ('pointwise' detection) are fundamentally limited in their ability to prevent fine-tuning attacks. We construct 'pointwise-undetectable' attacks that repurpose entropy in benign model outputs (e.g. semantic or syntactic variations) to covertly transmit dangerous knowledge. Our attacks are composed solely of unsuspicious benign samples that can be collected from the model before fine-tuning, meaning training and inference samples are all individually benign and low-perplexity. We test our attacks against the OpenAI fine-tuning API, finding they succeed in eliciting answers to harmful multiple-choice questions, and that they evade an enhanced monitoring system we design that successfully detects other fine-tuning attacks. We encourage the community to develop defences that tackle the fundamental limitations we uncover in pointwise fine-tuning API defences.


Human Misperception of Generative-AI Alignment: A Laboratory Experiment

arXiv.org Artificial Intelligence

We conduct an incentivized laboratory experiment to study people's perception of generative artificial intelligence (GenAI) alignment in the context of economic decision-making. Using a panel of economic problems spanning the domains of risk, time preference, social preference, and strategic interactions, we ask human subjects to make choices for themselves and to predict the choices made by GenAI on behalf of a human user. We find that people overestimate the degree of alignment between GenAI's choices and human choices. In every problem, human subjects' average prediction about GenAI's choice is substantially closer to the average human-subject choice than it is to the GenAI choice. At the individual level, different subjects' predictions about GenAI's choice in a given problem are highly correlated with their own choices in the same problem. We explore the implications of people overestimating GenAI alignment in a simple theoretical model.


Xbox Pushes Ahead With Muse, a New Generative AI Model. Devs Say 'Nobody Will Want This'

WIRED

Microsoft is wading deeper into generative artificial intelligence for gaming with Muse, a new AI model announced today. The model, which was trained on Ninja Theory's multiplayer game Bleeding Edge, can help Xbox game developers build parts of games, Microsoft says. Muse can understand the physics and 3D environment inside a game and generate visuals and reactions to players' movements. Among the various use cases for Muse that Microsoft outlines in its announcement, perhaps the most intriguing involves game preservation. The company says Muse AI can study games from its vast back catalog of classic titles and optimize them for modern hardware.


ChatGPT will now combat bias with new measures put forth by OpenAI

FOX News

Fox News Correspondent, William La Jeunesse, joins'Fox News Sunday' to discuss the evolution of A.I. and the push lawmakers are making to regulate it. OpenAI has announced a set of new measures to combat bias within its suite of products, including ChatGPT. The artificial intelligence (AI) company recently unveiled an updated Model Spec, a document that defines how OpenAI wants its models to behave in ChatGPT and the OpenAI API. The company says this iteration of the Model Spec builds on the foundational version released last May. "I think with a tool as powerful as this, one where people can access all sorts of different information, if you really believe we're moving to artificial general intelligence (AGI) one day, you have to be willing to share how you're steering the model," Laurentia Romaniuk, who works on model behavior at OpenAI, told Fox News Digital.


Why I'm deeply sceptical about comparisons between humans and machines

New Scientist

Artificial intelligence has humans beat – at least when it comes to games like chess and Go, identifying the 3D structure of proteins, generating investment strategies…the list goes on and on. Some argue that models like ChatGPT are already at the threshold of human intelligence. OpenAI head Sam Altman even threw his unborn child under the bus, claiming "my kid is never gonna grow up being smarter than AI". The capabilities of modern AI are certainly impressive, but I am deeply sceptical about comparisons between humans and machines.


When AI Thinks It Will Lose, It Sometimes Cheats, Study Finds

TIME - Tech

Complex games like chess and Go have long been used to test AI models' capabilities. But while IBM's Deep Blue defeated reigning world chess champion Garry Kasparov in the 1990s by playing by the rules, today's advanced AI models like OpenAI's o1-preview are less scrupulous. When sensing defeat in a match against a skilled chess bot, they don't always concede, instead sometimes opting to cheat by hacking their opponent so that the bot automatically forfeits the game. That is the finding of a new study from Palisade Research, shared exclusively with TIME ahead of its publication on Feb. 19, which evaluated seven state-of-the-art AI models for their propensity to hack. While slightly older AI models like OpenAI's GPT-4o and Anthropic's Claude Sonnet 3.5 needed to be prompted by researchers to attempt such tricks, o1-preview and DeepSeek R1 pursued the exploit on their own, indicating that AI systems may develop deceptive or manipulative strategies without explicit instruction.


Microsoft wants to use generative AI tool to help make video games

New Scientist

An artificial intelligence model from Microsoft can recreate realistic video game footage that the company says could help designers make games, but experts are unconvinced that the tool will be useful for most game developers. Neural networks that can produce coherent and accurate footage from video games are not new. A recent Google-created AI generated a fully playable version of the classic computer game Doom without access to the underlying game engine. The original Doom, however, was released in 1993; more modern games are far more complex, with sophisticated physics and computationally intensive graphics, which have proved trickier for AIs to faithfully recreate. Google creates self-replicating life from digital'primordial soup' Now, Katja Hofmann at Microsoft Research and her colleagues have developed an AI model called Muse, which can recreate full sequences of the multiplayer online battle game Bleeding Edge. These sequences appear to obey the game's underlying physics and keep players and in-game objects consistent over time, which implies that the model has grasped a deep understanding of the game, says Hofmann.


EU accused of leaving 'devastating' copyright loophole in AI Act

The Guardian

"What I do not understand is that we are supporting big tech instead of protecting European creative ideas and content." The EU's AI Act, which came into force last year, was already in the works when ChatGPT, an AI chatbot that can generate essays, jokes and job applications, burst into public consciousness in late 2022, becoming the fastest-growing consumer application in history. ChatGPT was developed by OpenAI, which is also behind the AI image generator Dall-E. He would like legislation to fill that gap, but said it would take years, after the European Commission's decision last week to withdraw the proposed AI Liability Act. "It might be getting very difficult.


Before Going to Tokyo, I Tried Learning Japanese With ChatGPT

WIRED

On the final day of my visit to Japan, I'm alone and floating in some skyscraper's rooftop hot springs, praying no one joins me. For the last few months, I've been using ChatGPT's Advanced Voice Mode as an AI language tutor, part of a test to judge generative AI's potential as both a learning tool and a travel companion. The excessive talking to both strangers and a chatbot on my phone was illuminating as well as exhausting. I'm ready to shut my yapper for a minute and enjoy the silence. When OpenAI launched ChatGPT late in 2022, it set off a firestorm of generative AI competition and public interest.


Local Differences, Global Lessons: Insights from Organisation Policies for International Legislation

arXiv.org Artificial Intelligence

The rapid adoption of AI across diverse domains has led to the development of organisational guidelines that vary significantly, even within the same sector. This paper examines AI policies in two domains, news organisations and universities, to understand how bottom-up governance approaches shape AI usage and oversight. By analysing these policies, we identify key areas of convergence and divergence in how organisations address risks such as bias, privacy, misinformation, and accountability. We then explore the implications of these findings for international AI legislation, particularly the EU AI Act, highlighting gaps where practical policy insights could inform regulatory refinements. Our analysis reveals that organisational policies often address issues such as AI literacy, disclosure practices, and environmental impact, areas that are underdeveloped in existing international frameworks. We argue that lessons from domain-specific AI policies can contribute to more adaptive and effective AI governance at the global level. This study provides actionable recommendations for policymakers seeking to bridge the gap between local AI practices and international regulations.