kaczynski
Derail Yourself: Multi-turn LLM Jailbreak Attack through Self-discovered Clues
Ren, Qibing, Li, Hao, Liu, Dongrui, Xie, Zhanxu, Lu, Xiaoya, Qiao, Yu, Sha, Lei, Yan, Junchi, Ma, Lizhuang, Shao, Jing
This study exposes the safety vulnerabilities of Large Language Models (LLMs) in multi-turn interactions, where malicious users can obscure harmful intents across several queries. We introduce ActorAttack, a novel multi-turn attack method inspired by actor-network theory, which models a network of semantically linked actors as attack clues to generate diverse and effective attack paths toward harmful targets. ActorAttack addresses two main challenges in multi-turn attacks: (1) concealing harmful intents by creating an innocuous conversation topic about the actor, and (2) uncovering diverse attack paths towards the same harmful target by leveraging LLMs' knowledge to specify the correlated actors as various attack clues. In this way, ActorAttack outperforms existing single-turn and multi-turn attack methods across advanced aligned LLMs, even for GPT-o1. We will publish a dataset called SafeMTData, which includes multi-turn adversarial prompts and safety alignment data, generated by ActorAttack. We demonstrate that models safety-tuned using our safety dataset are more robust to multi-turn attacks. Code is available at https://github.com/renqibing/ActorAttack.
Chapter 11: The AI Story
Computer Science is no more about computers than astronomy is about telescopes. When looms weave by themselves, man's slavery will end. Within thirty years, we will have the technological means to create superhuman intelligence. Shortly after, the human era will be ended. Today we are entirely dependent on machines.
Will 2018 be the year of the neo-luddite?
One of the great paradoxes of digital life – understood and exploited by the tech giants – is that we never do what we say. Poll after poll in the past few years has found that people are worried about online privacy and do not trust big tech firms with their data. But they carry on clicking and sharing and posting, preferring speed and convenience above all else. Last year was Silicon Valley's annus horribilis: a year of bots, Russian meddling, sexism, monopolistic practice and tax-minimising. But I think 2018 might be worse still: the year of the neo-luddite, when anti-tech words turn into deeds. The caricature of machine-wrecking mobs doesn't capture our new approach to tech.
The Unabomber: uncanny prophecies of a dangerous man
He predicted that machines would eventually displace people in the workplace and that this would ultimately put the human race at the mercy of technology. This was written on a typewriter at a time when the internet was in its infancy, desktop computers were large, boxy affairs too expensive for most of us, and artificial intelligence was a fringe science, treated with derision by most. He explained: "As society and the problems that face it become more and more complex and as machines become more and more intelligent, people will let machines make more and more of their decisions for them, simply because machine-made decisions will bring better results than man-made ones. "Eventually a stage may be reached at which the decisions necessary to keep the system running will be so complex that human beings will be incapable of making them intelligently. At that stage the machines will be in effective control. People won't be able to just turn the machines off, because they will be so dependent on them that turning them off would amount to suicide.
Why the Future Doesn't Need Us
Our most powerful 21st-century technologies – robotics, genetic engineering, and nanotech – are threatening to make humans an endangered species. From the moment I became involved in the creation of new technologies, their ethical dimensions have concerned me, but it was only in the autumn of 1998 that I became anxiously aware of how great are the dangers facing us in the 21st century. I can date the onset of my unease to the day I met Ray Kurzweil, the deservedly famous inventor of the first reading machine for the blind and many other amazing things. This article has been reproduced in a new format and may be missing content or contain faulty links. Contact wiredlabs@wired.com to report an issue. Ray and I were both speakers at George Gilder's Telecosm conference, and I encountered him by chance in the bar of the hotel after both our sessions were over. I was sitting with John Searle, a Berkeley philosopher who studies consciousness. While we were talking, Ray approached and a ...