Refined Direct Preference Optimization with Synthetic Data for Behavioral Alignment of LLMs
–arXiv.org Artificial Intelligence
In this paper, we introduce refined Direct Preference Optimization (rDPO), a method for improving the behavioral alignment of Large Language Models (LLMs) without the need for human-annotated data. The method involves creating synthetic data using self-critique prompting by a teacher LLM and then utilising an generalized DPO loss function to distil to a student LLM. The loss function incorporates an additional external reward model to improve the quality of synthetic data, making rDPO robust to potential noise in the synthetic dataset. Progress in large language models (LLMs) has broadened their application scope, but worries about their safe and ethical utilization continue to exist. A notable breakthrough in LLMs involves the posttraining alignment to desired behaviors (Chung et al., 2022). However, this process often depends on expensive human-annotated data. Common alignment strategies feature Supervised Fine-Tuning (SFT) (Tunstall et al., 2023) and Reinforcement Learning from Human Feedback (RLHF) (Ziegler et al., 2019; Christiano et al., 2017; Ouyang et al., 2022). Both methodologies heavily depend on extensive human annotation. Therefore, the community aims to develop fine-tuning strategies that can effectively leverage synthetic data, that is, data generated by an LLM, ultimately facilitating the alignment process. Our research aligns with the larger ambition of evolving weak models to strong ones, a fundamental concept in machine learning, rooted in distillation approaches (Hinton et al., 2015) that do not need extra annotated data.
arXiv.org Artificial Intelligence
Feb-12-2024
- Genre:
- Research Report (1.00)
- Industry:
- Information Technology > Security & Privacy (0.68)
- Health & Medicine (0.46)
- Banking & Finance (0.46)
- Technology: