Language Models that Think, Chat Better
Bhaskar, Adithya, Ye, Xi, Chen, Danqi
–arXiv.org Artificial Intelligence
Reinforcement learning with verifiable rewards (RL VR) improves language model reasoning by using rule-based rewards in verifiable domains such as mathematics and code. However, RL VR leads to limited generalization for open-ended tasks--such as writing outline essays or making meal plans--where humans reason routinely. This paper shows that the RL VR paradigm is effective beyond verifiable domains, and introduces RL with Model-rewarded Thinking (RLMT) for general-purpose chat capabilities. Using diverse real-world prompts, RLMT requires LMs to generate long CoT reasoning before response, and optimizes them with online RL against a preference-based reward model used in RLHF. Across 40 training runs on Llama-3.1-8B and Qwen-2.5-7B This includes substantial gains of 3-7 points on three chat benchmarks (AlpacaEval2, WildBench, and Arena-HardV2), along with 1-3 point improvements on other tasks like creative writing and general knowledge. Our best 8B model surpasses GPT -4o in chat and creative writing and rivals Claude-3.7-Sonnet RLMT can also be applied directly to base models without an SFT stage, akin to R1-Zero training (DeepSeek-AI, 2025). Remarkably, with only 7K prompts, Llama-3.1-8B We close with qualitative and quantitative analyses of how trained models plan their responses. Our results rethink the post-training pipeline and call upon future work to understand and employ thinking more broadly. HINKING through the consequences of one's actions--and revising them when needed--is a defining feature of human intelligence (often called "system 2 thinking", Kahneman (2011)). It has also become a central aspiration for large language models (LLMs).
arXiv.org Artificial Intelligence
Sep-25-2025
- Country:
- North America > United States > Minnesota (0.28)
- Genre:
- Research Report > New Finding (0.66)
- Industry:
- Health & Medicine (0.34)
- Technology: