LSPO: Length-aware Dynamic Sampling for Policy Optimization in LLM Reasoning

Chen, Weizhe, Koenig, Sven, Dilkina, Bistra

arXiv.org Artificial Intelligence 

Since the release of Deepseek-R1, reinforcement learning with verifiable rewards (RL VR) has become a central approach for training large language models (LLMs) on reasoning tasks. Recent work has largely focused on modifying loss functions to make RL VR more efficient and effective. In this paper, motivated by studies of overthinking in LLMs, we propose Length-aware Sampling for Policy Optimization (LSPO), a novel meta-RL VR algorithm that dynamically selects training data at each step based on the average response length. We evaluate LSPO across multiple base models and datasets, demonstrating that it consistently improves learning effectiveness. In addition, we conduct a detailed ablation study to examine alternative ways of incorporating length signals into dynamic sampling, offering further insights and highlighting promising directions for future research. Since the release of ChatGPT, large language models (LLMs) have rapidly evolved beyond traditional natural language tasks such as summarization and question answering, extending into broader domains of reasoning and problem solving Chen et al. (2023); Y ao et al. (2023); Chen et al. (2024a;b); Jaech et al. (2024); Huang & Y ang (2025). A growing body of research has therefore focused on how to make LLMs more effective and efficient on reasoning-intensive tasks. While a prominent line of work have investigated how to improve test-time performance without retraining a model from scratch Wang et al. (2022); Chen et al. (2023); Madaan et al. (2023); Gou et al. (2023); Guan et al. (2025); Chen et al. (2025), more recently, following the release of Deepseek-R1 DeepSeek-AI et al. (2025), reinforcement learning with verifiable rewards (RL VR) has emerged as a powerful and promising paradigm for post-training LLMs, yielding significant gains in reasoning ability. While Deepseek-R1 adopted GRPO Shao et al. (2024), subsequent research has proposed a variety of alternative algorithms to further enhance reasoning capabilities. In general, efforts to improve RL VR training have explored several dimensions, including alternative loss functions, improved datasets, and better hyperparameter tuning.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found