DogeRM: Equipping Reward Models with Domain Knowledge through Model Merging
Lin, Tzu-Han, Li, Chen-An, Lee, Hung-yi, Chen, Yun-Nung
–arXiv.org Artificial Intelligence
Reinforcement learning from human feedback (RLHF) is a popular strategy for aligning large language models (LLMs) with desired behaviors. Reward modeling is a crucial step in RLHF. However, collecting paired preference data for training reward models is often costly and time-consuming, especially for domain-specific preferences requiring expert annotation. To address this challenge, we propose the \textbf{Do}main knowled\textbf{ge} merged \textbf{R}eward \textbf{M}odel (DogeRM), a novel framework that integrates domain-specific knowledge into a general reward model by model merging. The experiments demonstrate that DogeRM enhances performance across different benchmarks and provide a detailed analysis showcasing the effects of model merging, showing the great potential of facilitating model alignment.
arXiv.org Artificial Intelligence
Jul-1-2024
- Country:
- North America
- United States > Virginia (0.04)
- Canada > Ontario
- Toronto (0.04)
- Europe
- Middle East > Malta
- Eastern Region > Northern Harbour District > St. Julian's (0.04)
- Ireland > Leinster
- County Dublin > Dublin (0.04)
- Middle East > Malta
- Asia > Taiwan
- Taiwan Province > Taipei (0.04)
- North America
- Genre:
- Research Report (1.00)
- Technology: