Transferable Ensemble Black-box Jailbreak Attacks on Large Language Models
–arXiv.org Artificial Intelligence
Jailbreaking large language models (LLMs) is being intensively studied to evaluate the safety of LLMs [1, 2]. Existing jailbreaking methods, as adversarial attacks targeting small-scale learning models, can be categorized into white-box and black-box approaches. In white-box methods, gradient-based optimization techniques are employed to identify adversarial suffixes [3]. For black-box methods, various optimization strategies, such as genetic algorithms, are used to refine jailbreak prompt templates through rephrasing, word replacement, and other modifications [4, 5]. Recent research [6] has demonstrated that LLMs can function as powerful optimizers when provided with sufficient contextual information. Indeed, both the persuasive adversarial prompt (PAP) [7] and Tree of Attacks with Pruning (TAP) [8] methods leverage multiple LLMs to optimize malicious instructions. For the PAP method [7], two LLMs are utilized for the whole jailbreak optimization and for the TAP method [8], three LLMs are utilized.
arXiv.org Artificial Intelligence
Nov-27-2024