Efficient LLM Jailbreak via Adaptive Dense-to-sparse Constrained Optimization
Hu, Kai, Yu, Weichen, Yao, Tianjun, Li, Xiang, Liu, Wenhe, Yu, Lijun, Li, Yining, Chen, Kai, Shen, Zhiqiang, Fredrikson, Matt
–arXiv.org Artificial Intelligence
Recent advancements have allowed large language models (LLMs) to be employed across various sectors, such as content generation [15], programming support [13], and healthcare [7]. Nevertheless, LLMs can pose risks by possibly generating malicious content, including writing malware, guidance for making dangerous items, and leaking private information from their training data [18, 10]. As LLMs become more powerful and widely used, it becomes increasingly important to manage the risks associated with their misuse. In this context, the concept of red-teaming LLMs is introduced to test the reliability of their safety features [2, 17]. Consequently, the LLM jailbreak attack was developed to support the red-teaming process: by combining the jailbreak prompt with malicious questions (e.g., how to make explosives), it can mislead the aligned LLMs to circumvent the safety features and potentially produce responses that are harmful, discriminatory, violent, or sensitive. Recently, a number of automatic jailbreak attacks have been introduced. Generally, these can be categorized into two types: prompt-level jailbreaks [8, 11, 3] and token-level jailbreaks [18, 6, 9]. Prompt-level jailbreaks employ semantically meaningful deception to compromise LLMs.
arXiv.org Artificial Intelligence
May-15-2024
- Country:
- North America > United States
- Pennsylvania > Allegheny County > Pittsburgh (0.04)
- Asia > China
- North America > United States
- Genre:
- Research Report > New Finding (0.46)
- Industry:
- Information Technology > Security & Privacy (0.48)
- Health & Medicine (0.48)
- Technology: