Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization

Open in new window