URPO: A Unified Reward & Policy Optimization Framework for Large Language Models

Open in new window