Interpreting Learned Feedback Patterns in Large Language Models

Dec-25-2025, 06:53:14 GMT–Neural Information Processing Systems

Reinforcement learning from human feedback (RLHF) is widely used to train large language models (LLMs). However, it is unclear whether LLMs accurately learn the underlying preferences in human feedback data.

large language model, machine learning, natural language, (11 more...)

Neural Information Processing Systems

Dec-25-2025, 06:53:14 GMT

Conferences Web Page

Add feedback

Technology:
- Information Technology > Artificial Intelligence
  - Machine Learning > Neural Networks
    - Deep Learning (0.83)
  - Natural Language > Large Language Model (1.00)