Adaptive Group Policy Optimization: Towards Stable Training and Token-Efficient Reasoning

Open in new window