Evaluating the Performance of Large Language Models for Spanish Language in Undergraduate Admissions Exams
Miranda, Sabino, Pichardo-Lagunas, Obdulia, Martínez-Seis, Bella, Baldi, Pierre
–arXiv.org Artificial Intelligence
This study evaluates the performance of large language models, specifically GPT-3.5 and BARD (supported by Gemini Pro model), in undergraduate admissions exams proposed by the National Polytechnic Institute in Mexico. The exams cover Engineering/Mathematical and Physical Sciences, Biological and Medical Sciences, and Social and Administrative Sciences. Both models demonstrated proficiency, exceeding the minimum acceptance scores for respective academic programs to up to 75% for some academic programs. GPT-3.5 outperformed BARD in Mathematics and Physics, while BARD performed better in History and questions related to factual information. Overall, GPT-3.5 marginally surpassed BARD with scores of 60.94% and 60.42%, respectively.
arXiv.org Artificial Intelligence
Dec-28-2023
- Country:
- North America
- United States > California
- Orange County > Irvine (0.14)
- Mexico > Mexico City
- Mexico City (0.04)
- United States > California
- Europe > Spain
- North America
- Genre:
- Research Report (0.50)
- Industry:
- Health & Medicine (1.00)
- Education > Educational Setting
- Higher Education (0.68)
- Technology: