GAPO: Learning Preferential Prompt through Generative Adversarial Policy Optimization

Open in new window