Bilingual Evaluation of Language Models on General Knowledge in University Entrance Exams with Minimal Contamination

Salido, Eva Sánchez, Morante, Roser, Gonzalo, Julio, Marco, Guillermo, Carrillo-de-Albornoz, Jorge, Plaza, Laura, Amigó, Enrique, Fernández, Andrés, Benito-Santos, Alejandro, Espinosa, Adrián Ghajari, Fresno, Victor

arXiv.org Artificial Intelligence 

The dataset contains 1003 multiple-choice questions in Spanish from various With the recent progress in broadening the generalisation subjects of the UNED Access Course for Over-25s capabilities of Large Language Models in Spanish, and high-quality English professional (LLMs), much current research focuses on understanding translations. Two characteristics make this dataset their capabilities and limitations. Evaluations unique: first, the contamination level should be of the most recent generative models, such minimal because UNED typically does not release as Llama-2 (Touvron et al., 2023b), Mistral (Jiang the answers to the exam questions, which are only et al., 2023), Mixtral (Jiang et al., 2024a), Gemini accessible to the teachers of each course. Second, (Anil et al., 2024), Gemma (Mesnard et al., this is a high-quality bilingual dataset, with original 2024), GPT-3.5 (Brown et al., 2020), GPT-4 and questions in Spanish translated into English GPT-4o (Achiam et al., 2024), attempt at measuring manually by a professional translators who did not their world knowledge, memorisation and inference use any external software.