Language Models that Think, Chat Better

Bhaskar, Adithya, Ye, Xi, Chen, Danqi

arXiv.org Artificial Intelligence 

Reinforcement learning with verifiable rewards (RL VR) improves language model reasoning by using rule-based rewards in verifiable domains such as mathematics and code. However, RL VR leads to limited generalization for open-ended tasks--such as writing outline essays or making meal plans--where humans reason routinely. This paper shows that the RL VR paradigm is effective beyond verifiable domains, and introduces RL with Model-rewarded Thinking (RLMT) for general-purpose chat capabilities. Using diverse real-world prompts, RLMT requires LMs to generate long CoT reasoning before response, and optimizes them with online RL against a preference-based reward model used in RLHF. Across 40 training runs on Llama-3.1-8B and Qwen-2.5-7B This includes substantial gains of 3-7 points on three chat benchmarks (AlpacaEval2, WildBench, and Arena-HardV2), along with 1-3 point improvements on other tasks like creative writing and general knowledge. Our best 8B model surpasses GPT -4o in chat and creative writing and rivals Claude-3.7-Sonnet RLMT can also be applied directly to base models without an SFT stage, akin to R1-Zero training (DeepSeek-AI, 2025). Remarkably, with only 7K prompts, Llama-3.1-8B We close with qualitative and quantitative analyses of how trained models plan their responses. Our results rethink the post-training pipeline and call upon future work to understand and employ thinking more broadly. HINKING through the consequences of one's actions--and revising them when needed--is a defining feature of human intelligence (often called "system 2 thinking", Kahneman (2011)). It has also become a central aspiration for large language models (LLMs).

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found