RLSR: Reinforcement Learning with Supervised Reward Outperforms SFT in Instruction Following

Open in new window