Group Policy Gradient
Chen, Junhua, Zhang, Zixi, Zhong, Hantao, Antonova, Rika
We introduce Group Policy Gradient (GPG), a family of critic-free policy-gradient estimators for general MDPs. Inspired by the success of GRPO's approach in Reinforcement Learning from Human Feedback (RLHF), GPG replaces a learned value function with a group-based Monte Carlo advantage estimator, removing the memory, compute, and hyperparameter costs of training a critic while preserving PPO's clipped-objective structure. We prove the consistency of the GPG estimator, analyze the bias-variance tradeoffs, and demonstrate empirically that GPG matches or outperforms PPO on standard benchmarks. GPG makes better use of parallel simulations, which, together with its critic-free design, results in more efficient use of computational resources than PPO.
Oct-7-2025
- Country:
- Europe
- United Kingdom > England
- Cambridgeshire > Cambridge (0.15)
- Portugal > Braga
- Braga (0.04)
- United Kingdom > England
- Asia > Middle East
- Jordan (0.04)
- Europe
- Genre:
- Research Report (0.52)
- Technology: