DAST: Difficulty-Aware Self-Training on Large Language Models
Xue, Boyang, Zhu, Qi, Wang, Hongru, Wang, Rui, Wang, Sheng, Xu, Hongling, Mi, Fei, Wang, Yasheng, Shang, Lifeng, Liu, Qun, Wong, Kam-Fai
–arXiv.org Artificial Intelligence
Present Large Language Models (LLM) self-training methods always under-sample on challenging queries, leading to inadequate learning on difficult problems which limits LLMs' ability. Therefore, this work proposes a difficulty-aware self-training (DAST) framework that focuses on improving both the quantity and quality of self-generated responses on challenging queries during self-training. DAST is specified in three components: 1) sampling-based difficulty level estimation, 2) difficulty-aware data augmentation, and 3) the self-training algorithm using SFT and DPO respectively. Experiments on mathematical tasks demonstrate the effectiveness and generalization of DAST, highlighting the critical role of difficulty-aware strategies in advancing LLM self-training.
arXiv.org Artificial Intelligence
Mar-11-2025
- Country:
- North America
- Mexico > Mexico City
- Mexico City (0.04)
- Canada > Ontario
- Toronto (0.04)
- Mexico > Mexico City
- Europe > Italy
- Calabria > Catanzaro Province > Catanzaro (0.04)
- Asia
- Singapore (0.04)
- Thailand > Bangkok
- Bangkok (0.04)
- Japan > Honshū
- Tōhoku > Iwate Prefecture > Morioka (0.04)
- China
- Hong Kong (0.04)
- Heilongjiang Province > Harbin (0.04)
- Guangdong Province > Shenzhen (0.04)
- North America
- Genre:
- Research Report > New Finding (0.46)
- Technology: