Goto

Collaborating Authors

 Large Language Model


What Features in Prompts Jailbreak LLMs? Investigating the Mechanisms Behind Attacks

arXiv.org Artificial Intelligence

While `jailbreaks' have been central to research on the safety and reliability of LLMs (large language models), the underlying mechanisms behind these attacks are not well understood. Some prior works have used linear methods to analyze jailbreak prompts or model refusal. Here, however, we compare linear and nonlinear methods to study the features in prompts that contribute to successful jailbreaks. We do this by probing for jailbreak success based only on the portions of the latent representations corresponding to prompt tokens. First, we introduce a dataset of 10,800 jailbreak attempts from 35 attack methods. We then show that different jailbreaking methods work via different nonlinear features in prompts. Specifically, we find that while probes can distinguish between successful and unsuccessful jailbreaking prompts with a high degree of accuracy, they often transfer poorly to held-out attack methods. We also show that nonlinear probes can be used to mechanistically jailbreak the LLM by guiding the design of adversarial latent perturbations. These mechanistic jailbreaks are able to jailbreak Gemma-7B-IT more reliably than 34 of the 35 techniques that it was trained on. Ultimately, our results suggest that jailbreaks cannot be thoroughly understood in terms of universal or linear prompt features alone.


Can Humans Oversee Agents to Prevent Privacy Leakage? A Study on Privacy Awareness, Preferences, and Trust in Language Model Agents

arXiv.org Artificial Intelligence

Language model (LM) agents that act on users' behalf for personal tasks can boost productivity, but are also susceptible to unintended privacy leakage risks. We present the first study on people's capacity to oversee the privacy implications of the LM agents. By conducting a task-based survey (N=300), we investigate how people react to and assess the response generated by LM agents for asynchronous interpersonal communication tasks, compared with a response they wrote. We found that people may favor the agent response with more privacy leakage over the response they drafted or consider both good, leading to an increased harmful disclosure from 15.7% to 55.0%. We further uncovered distinct patterns of privacy behaviors, attitudes, and preferences, and the nuanced interactions between privacy considerations and other factors. Our findings shed light on designing agentic systems that enable privacy-preserving interactions and achieve bidirectional alignment on privacy preferences to help users calibrate trust.


One Arrow, Many Targets: Probing LLMs for Multi-Attribute Controllable Text Summarization

arXiv.org Artificial Intelligence

Text summarization is a well-established task within the natural language processing (NLP) community. However, the focus on controllable summarization tailored to user requirements is gaining traction only recently. While several efforts explore controllability in text summarization, the investigation of Multi-Attribute Controllable Summarization (MACS) remains limited. This work addresses this gap by examining the MACS task through the lens of large language models (LLMs), using various learning paradigms, particularly low-rank adapters. We experiment with different popular adapter fine-tuning strategies to assess the effectiveness of the resulting models in retaining cues and patterns associated with multiple controllable attributes. Additionally, we propose and evaluate a novel hierarchical adapter fusion technique to integrate learnings from two distinct controllable attributes. Subsquently, we present our findings, discuss the challenges encountered, and suggest potential avenues for advancing the MACS task.


Reasoning Limitations of Multimodal Large Language Models. A case study of Bongard Problems

arXiv.org Artificial Intelligence

Abstract visual reasoning (AVR) encompasses a suite of tasks whose solving requires the ability to discover common concepts underlying the set of pictures through an analogy-making process, similarly to human IQ tests. Bongard Problems (BPs), proposed in 1968, constitute a fundamental challenge in this domain mainly due to their requirement to combine visual reasoning and verbal description. This work poses a question whether multimodal large language models (MLLMs) inherently designed to combine vision and language are capable of tackling BPs. To this end, we propose a set of diverse MLLM-suited strategies to tackle BPs and examine four popular proprietary MLLMs: GPT-4o, GPT-4 Turbo, Gemini 1.5 Pro, and Claude 3.5 Sonnet, and four open models: InternVL2-8B, LLaVa-1.6 Mistral-7B, Phi-3.5-Vision, and Pixtral 12B. The above MLLMs are compared on three BP datasets: a set of original BP instances relying on synthetic, geometry-based images and two recent datasets based on real-world images, i.e., Bongard-HOI and Bongard-OpenWorld. The experiments reveal significant limitations of MLLMs in solving BPs. In particular, the models struggle to solve the classical set of synthetic BPs, despite their visual simplicity. Though their performance ameliorates on real-world concepts expressed in Bongard-HOI and Bongard-OpenWorld, the models still have difficulty in utilizing new information to improve their predictions, as well as utilizing a dialog context window effectively. To capture the reasons of performance discrepancy between synthetic and real-world AVR domains, we propose Bongard-RWR, a new BP dataset consisting of real-world images that translates concepts from hand-crafted synthetic BPs to real-world concepts. The MLLMs' results on Bongard-RWR suggest that their poor performance on classical BPs is not due to domain specificity but rather reflects their general AVR limitations.


Amazon reportedly bumped back its AI-powered Alexa to next year

Engadget

If you're wondering what happened to Amazon's new and improved version of its Alexa voice assistant, you're not alone. Bloomberg reports that the new Alexa is still stuck in its developmental phase and Amazon has cut off access to its beta phase including its new "Let's Chat" phase. As a result, a planned late 2024 launch has been pushed back to next year. The problem seems to be with its large language models (LLMs). The new Alexa is designed to understand more complicated questions from users but it's also more likely to fail doing some of the most basic things the old version could do quite easily like create a timer or operate smart lights, according to a follow up report from The Verge. Amazon originally planned to unveil its new version of Alexa AI in October but now the timeline has been extended into next year.


The Next Big iOS Upgrade Is Going to Make Your iPhone Look Very, Very Strange

Slate

Apple Intelligence is here, and I look forward to everyone sharing the bonkers/pointless #AI summaries it now puts on your lock screen. "Pugsley is little fester" is one of my faves. On @washingtonpost, I also have "Harris to endorse Harris."


AI Will Understand Humans Better Than Humans Do

WIRED

Michal Kosinski is a Stanford research psychologist with a nose for timely subjects. He sees his work as not only advancing knowledge, but alerting the world to potential dangers ignited by the consequences of computer systems. His best-known projects involved analyzing the ways in which Facebook (now Meta) gained a shockingly deep understanding of its users from all the times they clicked "like" on the platform. Now he's shifted to the study of surprising things that AI can do. He's conducted experiments, for example, that indicate that computers could predict a person's sexuality by analyzing a digital photo of their face.


Engadget Podcast: Apple's M4 chip heads to the iMac, Mac mini and MacBook Pro

Engadget

It's been a Mac-heavy week! The Mac mini, in particular, looks like it'll be a huge hit for anyone who needs a simple desktop system. Also, we dive into why Apple is pushing for every Mac to get 16GB of RAM at a minimum. That will benefit all users, even if they don't care about Apple Intelligence. Listen below or subscribe on your podcast app of choice. If you've got suggestions or topics you'd like covered on the show, be sure to email us or drop a note in the comments! And be sure to check out our other podcast, Engadget News! Regulators force Lyft to tell U.S. drivers accurate numbers of how much money they'll make โ€“ 45:30 This week, I'm joined by our podcast producer, Ben Ellman. Kind of a light ship this week because a lot of people are out. Everyone's on taking some break and a lot of people are just busy at Engadget. So it's just going to be us. But we've got a lot of news to dive into all of Apple's new Macs with M4 chips, the M4 Pro and M4 Max as well, that they all just announced this week. There's a lot of new stuff and I'm excited to talk about it as always, folks. So if you're enjoying the show, please be sure to subscribe to us on iTunes or your podcatcher of choice. Leave us a review on iTunes. And also, yeah, you can join us Thursday mornings, typically around 1045 AM Eastern on our YouTube channel for our live stream so we can do some Q& A. In fact, we'll be including some of those questions and our answers later in this episode as well. Ben, you are somebody who I know is fully in the Mac ecosystem, and I also know you're very conscientious. Well, unfortunately, or for what you do, you're kind of there, but you're also very conscientious about how you upgrade, right? How did you feel about all these new Macs? Because we have the M4 iMac, we have an adorable new Mac mini, which is tiny, absolutely tiny, and M4 chips on the MacBook Pros. Is anything particularly compelling to you? Ben: So as I was reading up on the Mac, All of the stuff they released this week. That chip is four years old now. Ben: cut me like a knife. But that is M1 Classic, not M1 Pro. My research says that the M1 Pro is only two times slower than this new M4 Pro. Please fact check me on this. Send us an email at podcast adding gadget. If I didn't get that right. Devindra: I mean, you, you bring up a good point though, Ben, be sure to be very clear about what Apple is comparing its devices to, right? Because they often go back to base M1, which. Was released at the end of 2020 2020. It took a full year before we got the M4 Pro and M4 Max chips, right. Before they really expanded the line. Ben: you mean M1 Pro and M1 Max. So remember that there was that time difference when they, they just dropped the M1 on us and that was on the MacBook Air, MacBook Pro 13 inch, which was a fricking waste of time and the Mac mini, I believe back then, right.


One in 20 new Wikipedia pages seem to be written with the help of AI

New Scientist

Nearly 5 per cent of new Wikipedia pages that are published in English seem to contain text generated by artificial intelligence, which could reduce the site's reliability. Creston Brooks at Princeton University and his colleagues were wondering about the implications of recently rolled-out AI systems called large language models and how these may affect sources of information.


The Download: OpenAI launches search, and AI-generated video games

MIT Technology Review

The news: ChatGPT can now search the web for up-to-date answers to a user's queries. Previously it was restricted to generating answers from its training data, and had limited web search capabilities. But now, ChatGPT will automatically search the web in response to queries about recent information such as sports, stocks, or news of the day, and can deliver rich multi-media results. How to use it: The feature is available now for the chatbot's paying users, but OpenAI intends to make it available for free later, even when people are logged out. It also plans to combine search with its voice features.