On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization

Open in new window