Refined Direct Preference Optimization with Synthetic Data for Behavioral Alignment of LLMs

Open in new window