Enhancing LLM Safety via Constrained Direct Preference Optimization
Liu, Zixuan, Sun, Xiaolin, Zheng, Zizhan
–arXiv.org Artificial Intelligence
The rapidly increasing capabilities of large language models (LLMs) raise an urgent need to align AI systems with diverse human preferences to simultaneously enhance their usefulness and safety, despite the often conflicting nature of these goals. To address this important problem, a promising approach is to enforce a safety constraint at the fine-tuning stage through a constrained Reinforcement Learning from Human Feedback (RLHF) framework. This approach, however, is computationally expensive and often unstable. In this work, we introduce Constrained DPO (C-DPO), a novel extension of the recently proposed Direct Preference Optimization (DPO) approach for fine-tuning LLMs that is both efficient and lightweight. By integrating dual gradient descent and DPO, our method identifies a nearly optimal trade-off between helpfulness and harmlessness without using reinforcement learning. Empirically, our approach provides a safety guarantee to LLMs that is missing in DPO while achieving significantly higher rewards under the same safety constraint compared to a recently proposed safe RLHF approach. Warning: This paper contains example data that may be offensive or harmful. Large language models (LLMs) have demonstrated remarkable proficiency in tasks like chat completion, instruction following, coding, problem-solving, and decision-making. However, they also suffer from various weaknesses and vulnerabilities (Wang et al., 2023; Wei et al., 2023), which can be a barrier to their use in security and safety-critical domains. Techniques such as supervised finetuning (SFT) and reinforcement learning with human feedback (RLHF) or AI feedback (RLAIF) have been employed to align these models more closely with human preferences. Yet, they fall short in providing robust defense against strategically designed adversarial inputs. As noted in (Wei et al., 2023), the limitation stems from the conflicting objectives inherent in LLM training, such as helpfulness and harmlessness, which are challenging to balance using a single reward or preference model. A promising direction for enhancing safety is to decouple the reward and safety objectives, and fine-tune an LLM to optimize the expected reward subject to a safety constraint, with the objective and the constraint modeled using distinct datasets from human (or AI) feedback (Ji et al., 2023). By imposing a safety constraint, this approach can potentially lead to a safer model without diminishing its utility. Further, it fits naturally in the safe RL framework extensively studied recently (Gu et al., 2023). However, a direct application of safe RL techniques to LLM fine-tuning is unsatisfactory.
arXiv.org Artificial Intelligence
Mar-4-2024
- Country:
- North America > United States
- Virginia (0.04)
- Washington > King County
- Louisiana > Orleans Parish
- New Orleans (0.04)
- California
- San Francisco County > San Francisco (0.14)
- Los Angeles County > Pasadena (0.04)
- Santa Clara County
- Sunnyvale (0.04)
- Palo Alto (0.04)
- Mountain View (0.04)
- Cupertino (0.04)
- Europe > United Kingdom
- England > Cambridgeshire > Cambridge (0.04)
- North America > United States
- Genre:
- Research Report > New Finding (0.46)
- Industry:
- Information Technology (1.00)
- Technology: