Models like GPT -4 excel in medical question answering but may face challenges in the lack of interpretability when handling complex tasks in real clinical settings.
To address these challenges, we consider the two-step preference modeling procedure that first resolves the under-specification by selecting a context, and then evaluates preference with respect to the chosen context.