Interpreting Learned Feedback Patterns in Large Language Models
–Neural Information Processing Systems
Reinforcement learning from human feedback (RLHF) is widely used to train large language models (LLMs). However, it is unclear whether LLMs accurately learn the underlying preferences in human feedback data.
Neural Information Processing Systems
Dec-25-2025, 06:53:14 GMT
- Technology: