Education
Changing Answer Order Can Decrease MMLU Accuracy
Gupta, Vipul, Pantoja, David, Ross, Candace, Williams, Adina, Ung, Megan
For can affect multiple choice tests, for example, example, NLP model accuracy has been shown to when answers are presented in a different order be fairly brittle. For example, accuracy can drop during retest (Krosnick and Fabrigar, 1991; when researchers apply input alterations based Tellinghuisen and Sulikowski, 2008; Lions et al., on paraphrasing (Gan and Ng, 2019), word order 2022). However, as models do not have the biological changes (Gauthier and Levy, 2019; Ribeiro et al., limitations of humans, we may expect them 2020; Sinha et al., 2021a, 2022; Allen-Zhu and Li, to exhibit less variation than humans, or possibly 2023a,b; Berglund et al., 2023; Golovneva et al., even none at all. Thus, we claim that a model 2024; Kitouni et al., 2024), or other minor, largely should be robust to answer order changes: if it gets meaning-preserving input variations or perturbations the correct answer to a question when the answer (Belinkov and Bisk, 2018; Ebrahimi et al., is labeled'A', it should also always get the correct 2018; Jiang et al., 2020; Gao et al., 2021; Li et al., answer when it is labeled'C'. Put another way, 2021; Sinha et al., 2021b; Moradi and Samwald, the model should select the same answer for each 2021; Papakipos and Bitton, 2022; Qian et al., question, regardless of its label, for every possible 2022; Goodarzi et al., 2023; Sinha et al., 2023).
Learning Pareto Set for Multi-Objective Continuous Robot Control
Shu, Tianye, Shang, Ke, Gong, Cheng, Nan, Yang, Ishibuchi, Hisao
For a control problem with multiple conflicting objectives, there exists a set of Pareto-optimal policies called the Pareto set instead of a single optimal policy. When a multi-objective control problem is continuous and complex, traditional multi-objective reinforcement learning (MORL) algorithms search for many Pareto-optimal deep policies to approximate the Pareto set, which is quite resource-consuming. In this paper, we propose a simple and resource-efficient MORL algorithm that learns a continuous representation of the Pareto set in a high-dimensional policy parameter space using a single hypernet. The learned hypernet can directly generate various well-trained policy networks for different user preferences. We compare our method with two state-of-the-art MORL algorithms on seven multi-objective continuous robot control problems. Experimental results show that our method achieves the best overall performance with the least training parameters. An interesting observation is that the Pareto set is well approximated by a curved line or surface in a high-dimensional parameter space. This observation will provide insight for researchers to design new MORL algorithms.
Efficient Continual Pre-training by Mitigating the Stability Gap
Guo, Yiduo, Fu, Jie, Zhang, Huishuai, Zhao, Dongyan, Shen, Yikang
Continual pre-training has increasingly become the predominant approach for adapting Large Language Models (LLMs) to new domains. This process involves updating the pre-trained LLM with a corpus from a new domain, resulting in a shift in the training distribution. To study the behavior of LLMs during this shift, we measured the model's performance throughout the continual pre-training process. we observed a temporary performance drop at the beginning, followed by a recovery phase, a phenomenon known as the "stability gap," previously noted in vision models classifying new classes. To address this issue and enhance LLM performance within a fixed compute budget, we propose three effective strategies: (1) Continually pre-training the LLM on a subset with a proper size for multiple epochs, resulting in faster performance recovery than pre-training the LLM on a large corpus in a single epoch; (2) Pre-training the LLM only on high-quality sub-corpus, which rapidly boosts domain performance; and (3) Using a data mixture similar to the pre-training data to reduce distribution gap. We conduct various experiments on Llama-family models to validate the effectiveness of our strategies in both medical continual pre-training and instruction tuning. For example, our strategies improve the average medical task performance of the OpenLlama-3B model from 36.2% to 40.7% with only 40% of the original training budget and enhance the average general task performance without causing forgetting. Furthermore, we apply our strategies to the Llama-3-8B model. The resulting model, Llama-3-Physician, achieves the best medical performance among current open-source models, and performs comparably to or even better than GPT-4 on several medical benchmarks. We release our models at \url{https://huggingface.co/YiDuo1999/Llama-3-Physician-8B-Instruct}.
Number of girls in England taking computing GCSE plummets, study finds
The number of girls in England studying for a GCSE in computing has more than halved in less than a decade, prompting warnings about the "dominance of men in shaping the modern world". The sharp decline in female participation follows government qualification changes that led to the scrapping of the old information communication technology (ICT) GCSE and its replacement with a new computer science GCSE. While the government's reforms were aimed at creating "more academically challenging and knowledge-based" qualifications, the introduction of the new syllabus has had the unintended consequence of driving female entries down, according to new research by King's College London. In 2015 43% of candidates for ICT GCSE were female, compared with just 21% of those who took GCSE computer science in 2023. In numerical terms, 40,000 female students took ICT GCSE in 2015, with a further 5,000 taking computer science.
University examiners fail to spot ChatGPT answers in real-world test
Ninety-four per cent of university exam submissions created using ChatGPT weren't detected as being generated by artificial intelligence, and these submissions tended to get higher scores than real students' work. Peter Scarfe at the University of Reading, UK, and his colleagues used ChatGPT to produce answers to 63 assessment questions on five modules across the university's psychology undergraduate degrees. Students sat these exams at home, so they were allowed to look at notes and references, and they could potentially have used AI although this wasn't permitted. How this moment for AI will change society forever (and how it won't) The AI-generated answers were submitted alongside real students' work, and accounted for, on average, 5 per cent of the total scripts marked by academics. The markers weren't informed that they were checking the work of 33 fake students โ whose names were themselves generated by ChatGPT.
AI can beat university students, study suggests
In the study, fake exam answers and essays were submitted for first-, second- and third-year modules, without the knowledge of those marking them. The scores by the AI students beat those achieved by the real undergraduates in the first two years. But the humans scored better in the third-year exams - which "is consistent with the notion that current AI struggles with more abstract reasoning", the researchers said. And theirs was the largest and most robust blind study of its kind to date. Academics have raised concerns about the influence of AI in education, with Glasgow University recently reintroducing in-person exams for one course.
CheatGPT! Examiners struggle to tell the difference between answers written by AI and those from real human students - so, can you tell which of these papers was written by a bot?
The art of cheating in exams has come a long way since the days of scribbling a few notes on your wrist. In fact, a new study suggests AI chatbots are making cheating more efficient than ever. Even experienced examiners now struggle to spot the difference between AI and real human students, researchers have found. The experts from the University of Reading secretly added responses entirely generated by ChatGPT to a real undergraduate psychology exam. And, despite using AI in the simplest and most obvious manner, unsuspecting markers failed to spot the AI responses in 94 per cent of cases.
Researchers fool university markers with AI-generated exam papers
Researchers at the University of Reading fooled their own professors by secretly submitting AI-generated exam answers that went undetected and got better grades than real students. The project created fake student identities to submit unedited answers generated by ChatGPT-4 in take-home online assessments for undergraduate courses. The university's markers โ who were not told about the project โ flagged only one of the 33 entries, with the remaining AI answers receiving higher than average grades than the students. The authors said their findings showed that AI processors such as ChatGPT were now passing the "Turing test" โ named after the computing pioneer Alan Turing โ of being able to pass undetected by experienced judges. Billed as "the largest and most robust blind study of its kind" to investigate if human educators could detect AI-generated responses, the authors warned that it had major implications for how universities assess students. "Our research shows it is of international importance to understand how AI will affect the integrity of educational assessments," said Dr Peter Scarfe, one of the authors and an associate professor at Reading's school of psychology and clinical language sciences.
New computer vision method helps speed up screening of electronic materials
MIT graduate students Eunice Aissi, left, and Alexander Siemenn, have developed a technique that automatically analyzes visual features in printed samples (pictured) to quickly determine key properties of new and promising semiconducting materials. Boosting the performance of solar cells, transistors, LEDs, and batteries will require better electronic materials, made from novel compositions that have yet to be discovered. To speed up the search for advanced functional materials, scientists are using AI tools to identify promising materials from hundreds of millions of chemical formulations. In tandem, engineers are building machines that can print hundreds of material samples at a time based on chemical compositions tagged by AI search algorithms. But to date, there's been no similarly speedy way to confirm that these printed materials actually perform as expected.
Benchmarking General-Purpose In-Context Learning
Wang, Fan, Lin, Chuan, Cao, Yang, Kang, Yu
In-context learning (ICL) empowers generative models to address new tasks effectively and efficiently on the fly, without relying on any artificially crafted optimization techniques. In this paper, we study extending ICL to address a broader range of tasks with an extended learning horizon and higher improvement potential, namely General-Purpose In-Context Learning (GPICL). To this end, we introduce two lightweight benchmarks specifically crafted to train and evaluate GPICL functionalities. Each benchmark encompasses a vast number of tasks characterized by significant task variance, facilitating meta-training that minimizes inductive bias. These tasks are also crafted to promote long-horizon in-context learning through continuous generation and interaction. These characteristics necessitate the models to leverage contexts and history interactions to enhance their capabilities, across domains such as language modeling, decision-making, and world modeling. Our experiments on the baseline models demonstrate that meta-training with minimal inductive bias and ICL from the ground up is feasible across all the domains we've discussed. Additionally, our findings indicate that the scale of parameters alone may not be crucial for ICL or GPICL, suggesting alternative approaches such as increasing the scale of contexts and memory states.