Goto

Collaborating Authors

 Education


Key moments that defined education in America in 2023

FOX News

America's Newsroom anchor Bill Hemmer looks back at the top headlines of the past 12 months. Supreme Court rulings, wars waged over parental rights, crackdowns on conservative school boards and scandals that imbued some districts with controversy: Education's rocky landscape showcased this year's equally tumultuous cultural climate, and the issue has taken center stage for candidates going into 2024. Republicans continued to capitalize on parents' concerns that children are being exposed to age-inappropriate content in the classroom while calling for school choice and cautioning against giving transgender students access to single-sex spaces. Democrats, meanwhile, called out the opposition for alleged "book bans" and a majority defended transgender students' access to spaces corresponding with their preferred gender. The gridlock is expected to augment the intensity of an already explosive election season next year, and the issues aren't expected to fade anytime soon.


AI can benefit students and parents if done right

FOX News

Fox News medical contributor Dr. Marc Siegel outlines how the medical field is working to integrate artificial intelligence to care for heart conditions. If you've read the headlines in 2023, Artificial Intelligence (AI) is either coming to save or destroy education. From AI tools that can help proofread students' work to chatbots that can act as a kind of virtual research assistant, there are applications emerging that could rapidly improve what students are able to do and how they are able to do it. At the same time, ask any teacher, and you'll hear myriad stories of AI-generated essays (many with incorrect information in them), and the yeoman's work necessary to ChatGPT-proof their tests and quizzes. Rather than look at the whole AI and education universe, let's focus on one significant challenge confronting K-12 education today.


Exploring the Sensitivity of LLMs' Decision-Making Capabilities: Insights from Prompt Variation and Hyperparameters

arXiv.org Artificial Intelligence

The advancement of Large Language Models (LLMs) has led to their widespread use across a broad spectrum of tasks including decision making. Prior studies have compared the decision making abilities of LLMs with those of humans from a psychological perspective. However, these studies have not always properly accounted for the sensitivity of LLMs' behavior to hyperparameters and variations in the prompt. In this study, we examine LLMs' performance on the Horizon decision making task studied by Binz and Schulz (2023) analyzing how LLMs respond to variations in prompts and hyperparameters. By experimenting on three OpenAI language models possessing different capabilities, we observe that the decision making abilities fluctuate based on the input prompts and temperature settings. Contrary to previous findings language models display a human-like exploration exploitation tradeoff after simple adjustments to the prompt.


Automatic Essay Scoring in a Brazilian Scenario

arXiv.org Artificial Intelligence

The evolution of educational assessment methods has been influenced by technological advancements, particularly in the realm of Automatic Essay Scoring (AES)[1]. This technology, leveraging the power of artificial intelligence, has emerged as a promising tool in evaluating written responses, especially in large-scale settings. The adoption of AES has become increasingly pertinent in countries like Brazil, where standardized tests play a crucial role in determining access to higher education. However, the journey towards integrating AES in such contexts has unique challenges and considerations. One of the primary challenges faced in Brazil's educational assessment is the logistical and financial limitations associated with the traditional human grading system. With vast numbers of students participating in key examinations, the process of grading becomes not only time-consuming but also a significant financial burden on the educational system. This situation often leads to prolonged waiting periods. Recognizing these challenges, this research introduces an automatic grading algorithm specifically designed for Portuguese-language essays.


ReliCD: A Reliable Cognitive Diagnosis Framework with Confidence Awareness

arXiv.org Artificial Intelligence

During the past few decades, cognitive diagnostics modeling has attracted increasing attention in computational education communities, which is capable of quantifying the learning status and knowledge mastery levels of students. Indeed, the recent advances in neural networks have greatly enhanced the performance of traditional cognitive diagnosis models through learning the deep representations of students and exercises. Nevertheless, existing approaches often suffer from the issue of overconfidence in predicting students' mastery levels, which is primarily caused by the unavoidable noise and sparsity in realistic student-exercise interaction data, severely hindering the educational application of diagnostic feedback. To address this, in this paper, we propose a novel Reliable Cognitive Diagnosis(ReliCD) framework, which can quantify the confidence of the diagnosis feedback and is flexible for different cognitive diagnostic functions. Specifically, we first propose a Bayesian method to explicitly estimate the state uncertainty of different knowledge concepts for students, which enables the confidence quantification of diagnostic feedback. In particular, to account for potential differences, we suggest modeling individual prior distributions for the latent variables of different ability concepts using a pre-trained model. Additionally, we introduce a logical hypothesis for ranking confidence levels. Along this line, we design a novel calibration loss to optimize the confidence parameters by modeling the process of student performance prediction. Finally, extensive experiments on four real-world datasets clearly demonstrate the effectiveness of our ReliCD framework.


Online Algorithmic Recourse by Collective Action

arXiv.org Artificial Intelligence

Research on algorithmic recourse typically considers how an individual can reasonably change an unfavorable automated decision when interacting with a fixed decision-making system. This paper focuses instead on the online setting, where system parameters are updated dynamically according to interactions with data subjects. Beyond the typical individual-level recourse, the online setting opens up new ways for groups to shape system decisions by leveraging the parameter update rule. We show empirically that recourse can be improved when users coordinate by jointly computing their feature perturbations, underscoring the importance of collective action in mitigating adverse automated decisions.


ChatEd: A Chatbot Leveraging ChatGPT for an Enhanced Learning Experience in Higher Education

arXiv.org Artificial Intelligence

With the rapid evolution of Natural Language Processing (NLP), Large Language Models (LLMs) like ChatGPT have emerged as powerful tools capable of transforming various sectors. Their vast knowledge base and dynamic interaction capabilities represent significant potential in improving education by operating as a personalized assistant. However, the possibility of generating incorrect, biased, or unhelpful answers are a key challenge to resolve when deploying LLMs in an education context. This work introduces an innovative architecture that combines the strengths of ChatGPT with a traditional information retrieval based chatbot framework to offer enhanced student support in higher education. Our empirical evaluations underscore the high promise of this approach.


Adaptive Control Strategy for Quadruped Robots in Actuator Degradation Scenarios

arXiv.org Artificial Intelligence

Quadruped robots have strong adaptability to extreme environments but may also experience faults. Once these faults occur, robots must be repaired before returning to the task, reducing their practical feasibility. One prevalent concern among these faults is actuator degradation, stemming from factors like device aging or unexpected operational events. Traditionally, addressing this problem has relied heavily on intricate fault-tolerant design, which demands deep domain expertise from developers and lacks generalizability. Learning-based approaches offer effective ways to mitigate these limitations, but a research gap exists in effectively deploying such methods on real-world quadruped robots. This paper introduces a pioneering teacher-student framework rooted in reinforcement learning, named Actuator Degradation Adaptation Transformer (ADAPT), aimed at addressing this research gap. This framework produces a unified control strategy, enabling the robot to sustain its locomotion and perform tasks despite sudden joint actuator faults, relying exclusively on its internal sensors. Empirical evaluations on the Unitree A1 platform validate the deployability and effectiveness of Adapt on real-world quadruped robots, and affirm the robustness and practicality of our approach.


The Tyranny of Possibilities in the Design of Task-Oriented LLM Systems: A Scoping Survey

arXiv.org Artificial Intelligence

This scoping survey focuses on our current understanding of the design space for task-oriented LLM systems and elaborates on definitions and relationships among the available design parameters. The paper begins by defining a minimal task-oriented LLM system and exploring the design space of such systems through a thought experiment contemplating the performance of diverse LLM system configurations (involving single LLMs, single LLM-based agents, and multiple LLM-based agent systems) on a complex software development task and hypothesizes the results. We discuss a pattern in our results and formulate them into three conjectures. While these conjectures may be partly based on faulty assumptions, they provide a starting point for future research. The paper then surveys a select few design parameters: covering and organizing research in LLM augmentation, prompting techniques, and uncertainty estimation, and discussing their significance. The paper notes the lack of focus on computational and energy efficiency in evaluating research in these areas. Our survey findings provide a basis for developing the concept of linear and non-linear contexts, which we define and use to enable an agent-centric projection of prompting techniques providing a lens through which prompting techniques can be viewed as multi-agent systems. The paper discusses the implications of this lens, for the cross-pollination of research between LLM prompting and LLM-based multi-agent systems; and also, for the generation of synthetic training data based on existing prompting techniques in research. In all, the scoping survey presents seven conjectures that can help guide future research efforts.


Building Efficient Universal Classifiers with Natural Language Inference

arXiv.org Artificial Intelligence

Generative Large Language Models (LLMs) have become the mainstream choice for fewshot and zeroshot learning thanks to the universality of text generation. Many users, however, do not need the broad capabilities of generative LLMs when they only want to automate a classification task. Smaller BERT-like models can also learn universal tasks, which allow them to do any text classification task without requiring fine-tuning (zeroshot classification) or to learn new tasks with only a few examples (fewshot), while being significantly more efficient than generative LLMs. This paper (1) explains how Natural Language Inference (NLI) can be used as a universal classification task that follows similar principles as instruction fine-tuning of generative LLMs, (2) provides a step-by-step guide with reusable Jupyter notebooks for building a universal classifier, and (3) shares the resulting universal classifier that is trained on 33 datasets with 389 diverse classes. Parts of the code we share has been used to train our older zeroshot classifiers that have been downloaded more than 55 million times via the Hugging Face Hub as of December 2023. Our new classifier improves zeroshot performance by 9.4%.