When All Options Are Wrong: Evaluating Large Language Model Robustness with Incorrect Multiple-Choice Options
Góral, Gracjan, Wiśnios, Emilia
–arXiv.org Artificial Intelligence
This paper examines the zero-shot ability of Large Language Models (LLMs) to detect multiple-choice questions with no correct answer, a crucial aspect of educational assessment quality. We explore this ability not only as a measure of subject matter knowledge but also as an indicator of critical thinking within LLMs. Our experiments, utilizing a range of LLMs on diverse questions, highlight the significant performance gap between questions with a single correct answer and those without. Llama-3.1-405B stands out by successfully identifying the lack of a valid answer in many instances. These findings suggest that LLMs should prioritize critical thinking over blind instruction following and caution against their use in educational settings where questions with incorrect answers might lead to inaccurate evaluations. This research sets a benchmark for assessing critical thinking in LLMs and emphasizes the need for ongoing model alignment to ensure genuine user comprehension and assistance.
arXiv.org Artificial Intelligence
Aug-27-2024
- Country:
- North America
- United States > New York
- New York County > New York City (0.04)
- Mexico > Mexico City
- Mexico City (0.04)
- United States > New York
- Europe
- Poland > Masovia Province
- Warsaw (0.04)
- Middle East > Malta
- Eastern Region > Northern Harbour District > St. Julian's (0.04)
- Poland > Masovia Province
- Asia
- Singapore (0.04)
- Indonesia > Bali (0.04)
- British Indian Ocean Territory > Diego Garcia (0.04)
- Thailand > Bangkok
- Bangkok (0.04)
- Myanmar > Tanintharyi Region
- Dawei (0.04)
- Middle East
- Jordan (0.04)
- Israel (0.04)
- Saudi Arabia > Asir Province
- Abha (0.04)
- Japan > Honshū
- Chūbu > Toyama Prefecture > Toyama (0.04)
- Africa > Zambia
- Southern Province > Choma (0.04)
- North America
- Genre:
- Research Report > New Finding (1.00)
- Industry:
- Education
- Educational Technology > Educational Software (0.67)
- Educational Setting (0.48)
- Assessment & Standards (0.48)
- Education
- Technology: