Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning

Open in new window