Education
A Closer Look at the Self-Verification Abilities of Large Language Models in Logical Reasoning
Hong, Ruixin, Zhang, Hongming, Pang, Xinyu, Yu, Dong, Zhang, Changshui
Logical reasoning has been an ongoing pursuit in the field of AI. Despite significant advancements made by large language models (LLMs), they still struggle with complex logical reasoning problems. To enhance reasoning performance, one promising direction is scalable oversight, which requires LLMs to identify their own errors and then improve by themselves. Various self-verification methods have been proposed in pursuit of this goal. Nevertheless, whether existing models understand their own errors well is still under investigation. In this paper, we take a closer look at the self-verification abilities of LLMs in the context of logical reasoning, focusing on their ability to identify logical fallacies accurately. We introduce a dataset, FALLACIES, containing 232 types of reasoning fallacies categorized in a hierarchical taxonomy. By conducting exhaustive experiments on FALLACIES, we obtain comprehensive and detailed analyses of a series of models on their verification abilities. Our main findings suggest that existing LLMs could struggle to identify fallacious reasoning steps accurately and may fall short of guaranteeing the validity of self-verification methods. Drawing from these observations, we offer suggestions for future research and practical applications of self-verification methods.
Boarding for ISS: Imbalanced Self-Supervised: Discovery of a Scaled Autoencoder for Mixed Tabular Datasets
Stocksieker, Samuel, Pommeret, Denys, Charpentier, Arthur
The field of imbalanced self-supervised learning, especially in the context of tabular data, has not been extensively studied. Existing research has predominantly focused on image datasets. This paper aims to fill this gap by examining the specific challenges posed by data imbalance in self-supervised learning in the domain of tabular data, with a primary focus on autoencoders. Autoencoders are widely employed for learning and constructing a new representation of a dataset, particularly for dimensionality reduction. They are also often used for generative model learning, as seen in variational autoencoders. When dealing with mixed tabular data, qualitative variables are often encoded using a one-hot encoder with a standard loss function (MSE or Cross Entropy). In this paper, we analyze the drawbacks of this approach, especially when categorical variables are imbalanced. We propose a novel metric to balance learning: a Multi-Supervised Balanced MSE. This approach reduces the reconstruction error by balancing the influence of variables. Finally, we empirically demonstrate that this new metric, compared to the standard MSE: i) outperforms when the dataset is imbalanced, especially when the learning process is insufficient, and ii) provides similar results in the opposite case.
The Ethics of AI in Education
Porayska-Pomsta, Kaska, Holmes, Wayne, Nemorin, Selena
The advent of big data, and of Artificial Intelligence (AI) applications that collect and consume such data, has led to fundamental questions about the ethics of AI designs and to efforts aimed to highlight and safeguard against any potential harms caused by the deployment of AI across diverse domains of applications. Typically, questions raised relate to the trustworthiness of AI as agent technologies that autonomously or semi-autonomously operate in human environments and that have the ability to alter human behaviour. Other questions concern the role that AI may play now and in the future in either resolving or amplifying pre-existing social biases and any resulting harms. Specifically, Ethical AI as an emergent area of AI research and policy, has been spurred by the revelations of AI applications (usually unintentionally) promoting and amplifying many of the discriminatory and oppressive practices, and assumptions that underpin pre-existing social and institutional systems, e.g., historical biases against non-dominant populations, against users characterised by some divergence from the so-called cognitive or physical'norm', or those who are socio-economically disadvantaged (Crawford, 2017a; Madaio et al., 2022; Porayska-Pomsta and Rajendran, 2019; Williamson, Eynon, Knox & Davis, in this volume). Numerous examples of AI bias are both well-documented and rehearsed throughout the emergent ethics of AI literature, in hundreds of policy reports about AI ethics and governance that have been published to date (c.f.
Generative AI in Education: A Study of Educators' Awareness, Sentiments, and Influencing Factors
Ghimire, Aashish, Prather, James, Edwards, John
The rapid advancement of artificial intelligence (AI) and the expanding integration of large language models (LLMs) have ignited a debate about their application in education. This study delves into university instructors' experiences and attitudes toward AI language models, filling a gap in the literature by analyzing educators' perspectives on AI's role in the classroom and its potential impacts on teaching and learning. The objective of this research is to investigate the level of awareness, overall sentiment towardsadoption, and the factors influencing these attitudes for LLMs and generative AI-based tools in higher education. Data was collected through a survey using a Likert scale, which was complemented by follow-up interviews to gain a more nuanced understanding of the instructors' viewpoints. The collected data was processed using statistical and thematic analysis techniques. Our findings reveal that educators are increasingly aware of and generally positive towards these tools. We find no correlation between teaching style and attitude toward generative AI. Finally, while CS educators show far more confidence in their technical understanding of generative AI tools and more positivity towards them than educators in other fields, they show no more confidence in their ability to detect AI-generated work.
Early Period of Training Impacts Out-of-Distribution Generalization
Liu, Chen Cecilia, Gurevych, Iryna
Prior research has found that differences in the early period of neural network training significantly impact the performance of in-distribution (ID) tasks. However, neural networks are often sensitive to out-of-distribution (OOD) data, making them less reliable in downstream applications. Yet, the impact of the early training period on OOD generalization remains understudied due to its complexity and lack of effective analytical methodologies. In this work, we investigate the relationship between learning dynamics and OOD generalization during the early period of neural network training. We utilize the trace of Fisher Information and sharpness, with a focus on gradual unfreezing (i.e. progressively unfreezing parameters during training) as the methodology for investigation. Through a series of empirical experiments, we show that 1) selecting the number of trainable parameters at different times during training, i.e. realized by gradual unfreezing -- has a minuscule impact on ID results, but greatly affects the generalization to OOD data; 2) the absolute values of sharpness and trace of Fisher Information at the initial period of training are not indicative for OOD generalization, but the relative values could be; 3) the trace of Fisher Information and sharpness may be used as indicators for the removal of interventions during early period of training for better OOD generalization.
Improving Forward Compatibility in Class Incremental Learning by Increasing Representation Rank and Feature Richness
Kim, Jaeill, Lee, Wonseok, Eo, Moonjung, Rhee, Wonjong
Class Incremental Learning (CIL) constitutes a pivotal subfield within continual learning, aimed at enabling models to progressively learn new classification tasks while retaining knowledge obtained from prior tasks. Although previous studies have predominantly focused on backward compatible approaches to mitigate catastrophic forgetting, recent investigations have introduced forward compatible methods to enhance performance on novel tasks and complement existing backward compatible methods. In this study, we introduce an effective-Rank based Feature Richness enhancement (RFR) method, designed for improving forward compatibility. Specifically, this method increases the effective rank of representations during the base session, thereby facilitating the incorporation of more informative features pertinent to unseen novel tasks. Consequently, RFR achieves dual objectives in backward and forward compatibility: minimizing feature extractor modifications and enhancing novel task performance, respectively. To validate the efficacy of our approach, we establish a theoretical connection between effective rank and the Shannon entropy of representations. Subsequently, we conduct comprehensive experiments by integrating RFR into eleven well-known CIL methods. Our results demonstrate the effectiveness of our approach in enhancing novel-task performance while mitigating catastrophic forgetting. Furthermore, our method notably improves the average incremental accuracy across all eleven cases examined.
Introduction to Human-Robot Interaction: A Multi-Perspective Introductory Course
In this paper I describe the design of an introductory course in These course goals are critically conditioned on the expected background Human-Robot Interaction. This project-driven course is designed to of the enrolled students. The course is offered at a small introduce undergraduate and graduate engineering students, especially engineering-only university with a strong focus on Robotics related those enrolled in Computer Science, Mechanical Engineering, fields (50% of all undergraduate students are enrolled in Mechanical and Robotics degree programs, to key theories and methods used Engineering or Computer Science degree programs, and degree programs in the field of Human-Robot Interaction that they would otherwise offered in Robotics at both the undergraduate and graduate be unlikely to see in those degree programs. To achieve this aim, level), but with no degree programs offered in social sciences or humanities the course takes students all the way from stakeholder analysis (e.g., Psychology) and few, if any, elective courses available to empirical evaluation, covering and integrating key Qualitative, in those fields. The university size and focus means that the course Design, Computational, and Quantitative methods along the way. I is offered at a mixed undergraduate/graduate level, and is primarily detail the goals, audience, and format of the course, and provide a offered to students from Computer Science, Mechanical Engineering, detailed walkthrough of the course syllabus.
Continual Vision-and-Language Navigation
Jeong, Seongjun, Kang, Gi-Cheon, Choi, Seongho, Kim, Joochan, Zhang, Byoung-Tak
Vision-and-Language Navigation (VLN) agents navigate to a destination using natural language instructions and the visual information they observe. Existing methods for training VLN agents presuppose fixed datasets, leading to a significant limitation: the introduction of new environments necessitates retraining with previously encountered environments to preserve their knowledge. This makes it difficult to train VLN agents that operate in the ever-changing real world. To address this limitation, we present the Continual Vision-and-Language Navigation (CVLN) paradigm, designed to evaluate agents trained through a continual learning process. For the training and evaluation of CVLN agents, we re-arrange existing VLN datasets to propose two datasets: CVLN-I, focused on navigation via initial-instruction interpretation, and CVLN-D, aimed at navigation through dialogue with other agents. Furthermore, we propose two novel rehearsal-based methods for CVLN, Perplexity Replay (PerpR) and Episodic Self-Replay (ESR). PerpR prioritizes replaying challenging episodes based on action perplexity, while ESR replays previously predicted action logits to preserve learned behaviors. We demonstrate the effectiveness of the proposed methods on CVLN through extensive experiments.
LLM2LLM: Boosting LLMs with Novel Iterative Data Enhancement
Lee, Nicholas, Wattanawong, Thanakul, Kim, Sehoon, Mangalam, Karttikeya, Shen, Sheng, Anumanchipali, Gopala, Mahoney, Michael W., Keutzer, Kurt, Gholami, Amir
Pretrained large language models (LLMs) are currently state-of-the-art for solving the vast majority of natural language processing tasks. While many real-world applications still require fine-tuning to reach satisfactory levels of performance, many of them are in the low-data regime, making fine-tuning challenging. To address this, we propose LLM2LLM, a targeted and iterative data augmentation strategy that uses a teacher LLM to enhance a small seed dataset by augmenting additional data that can be used for fine-tuning on a specific task. LLM2LLM (1) fine-tunes a baseline student LLM on the initial seed data, (2) evaluates and extracts data points that the model gets wrong, and (3) uses a teacher LLM to generate synthetic data based on these incorrect data points, which are then added back into the training data. This approach amplifies the signal from incorrectly predicted data points by the LLM during training and reintegrates them into the dataset to focus on more challenging examples for the LLM. Our results show that LLM2LLM significantly enhances the performance of LLMs in the low-data regime, outperforming both traditional fine-tuning and other data augmentation baselines. LLM2LLM reduces the dependence on labor-intensive data curation and paves the way for more scalable and performant LLM solutions, allowing us to tackle data-constrained domains and tasks. We achieve improvements up to 24.2% on the GSM8K dataset, 32.6% on CaseHOLD, 32.0% on SNIPS, 52.6% on TREC and 39.8% on SST-2 over regular fine-tuning in the low-data regime using a LLaMA2-7B student model.
CoLLEGe: Concept Embedding Generation for Large Language Models
Teehan, Ryan, Lake, Brenden, Ren, Mengye
Current language models are unable to quickly learn new concepts on the fly, often requiring a more involved finetuning process to learn robustly. Prompting in-context is not robust to context distractions, and often fails to confer much information about the new concepts. Classic methods for few-shot word learning in NLP, relying on global word vectors, are less applicable to large language models. In this paper, we introduce a novel approach named CoLLEGe (Concept Learning with Language Embedding Generation) to modernize few-shot concept learning. CoLLEGe is a meta-learning framework capable of generating flexible embeddings for new concepts using a small number of example sentences or definitions. Our primary meta-learning objective is simply to facilitate a language model to make next word predictions in forthcoming sentences, making it compatible with language model pretraining. We design a series of tasks to test new concept learning in challenging real-world scenarios, including new word acquisition, definition inference, and verbal reasoning, and demonstrate that our method succeeds in each setting without task-specific training.