PGPO: Enhancing Agent Reasoning via Pseudocode-style Planning Guided Preference Optimization

Open in new window