S2J: Bridging the Gap Between Solving and Judging Ability in Generative Reward Models
Sun, Shaoning, Yu, Jiachen, Wang, Zongqi, Yang, Xuewei, Gu, Tianle, Yang, Yujiu
–arXiv.org Artificial Intelligence
With the rapid development of large language models (LLMs), generative reward models (GRMs) have been widely adopted for reward modeling and evaluation. Previous studies have primarily focused on training specialized GRMs by optimizing them on preference datasets with the judgment correctness as supervision. While it's widely accepted that GRMs with stronger problem-solving capabilities typically exhibit superior judgment abilities, we first identify a significant solve-to-judge gap when examining individual queries. Specifically, the solve-to-judge gap refers to the phenomenon where GRMs struggle to make correct judgments on some queries (14%-37%), despite being fully capable of solving them. In this paper, we propose the Solve-to-Judge (S2J) approach to address this problem. Our comprehensive experiments demonstrate that S2J effectively reduces the solve-to-judge gap by 16.2%, thereby enhancing the model's judgment performance by 5.8%. Notably, S2J achieves state-of-the-art (SOT A) performance among GRMs built on the same base model while utilizing a significantly smaller training dataset. Moreover, S2J accomplishes this through self-evolution without relying on more powerful external models for distillation. As Large Language Models (LLMs) continue to evolve rapidly, a variety of evaluation paradigms have been proposed to accurately evaluate the quality of their responses. This is not only crucial for providing accurate reward signals in post-training (Ouyang et al., 2022; Bai et al., 2022; Wang et al., 2024a), but also important for automated evaluation and benchmark construction (Zheng et al., 2023; Dubois et al., 2024). Among them, Generative Reward Mmodels (GRMs) have been proposed as a solution, which treats evaluation as a capability of LLMs and leverages LLMs to evaluate other LLMs (Zheng et al., 2023; Li et al., 2025). Unlike scalar reward models, which only output a single numerical score (Liu et al., 2024a; Lambert et al., 2024), GRMs utilize the generative capabilities of LLMs to produce an interpretable analysis before rendering a verdict.
arXiv.org Artificial Intelligence
Sep-29-2025