Education
A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks
Bullwinkel, Blake, Russinovich, Mark, Salem, Ahmed, Zanella-Beguelin, Santiago, Jones, Daniel, Severi, Giorgio, Kim, Eugenia, Hines, Keegan, Minnich, Amanda, Zunger, Yonatan, Kumar, Ram Shankar Siva
Recent research has demonstrated that state-of-the-art LLMs and defenses remain susceptible to multi-turn jailbreak attacks. These attacks require only closed-box model access and are often easy to perform manually, posing a significant threat to the safe and secure deployment of LLM-based systems. We study the effectiveness of the Crescendo multi-turn jailbreak at the level of intermediate model representations and find that safety-aligned LMs often represent Crescendo responses as more benign than harmful, especially as the number of conversation turns increases. Our analysis indicates that at each turn, Crescendo prompts tend to keep model outputs in a "benign" region of representation space, effectively tricking the model into fulfilling harmful requests. Further, our results help explain why single-turn jailbreak defenses like circuit breakers are generally ineffective against multi-turn attacks, motivating the development of mitigations that address this generalization gap.
Toward Cyclic A.I. Modelling of Self-Regulated Learning: A Case Study with E-Learning Trace Data
Schwabe, Andrew, Akgรผn, รzgรผr, Haig, Ella
Many e-learning platforms assert their ability or potential to improve students' self-regulated learning (SRL), however the cyclical and undirected nature of SRL theoretical models represent significant challenges for representation within contemporary machine learning frameworks. We apply SRL-informed features to trace data in order to advance modelling of students' SRL activities, to improve predictability and explainability regarding the causal effects of learning in an eLearning environment. We demonstrate that these features improve predictive accuracy and validate the value of further research into cyclic modelling techniques for SRL.
Casper: Inferring Diverse Intents for Assistive Teleoperation with Vision Language Models
Liu, Huihan, Shah, Rutav, Liu, Shuijing, Pittenger, Jack, Seo, Mingyo, Cui, Yuchen, Bisk, Yonatan, Martรญn-Martรญn, Roberto, Zhu, Yuke
Deploying robots in human-centric settings like households requires balancing robot autonomy with humans' sense of agency [1, 2, 3, 4, 5, 6]. Full teleoperation offers users fine-grained control but imposes a high cognitive load, whereas fully autonomous robots act independently but often misalign their actions with nuanced human needs. Assistive teleoperation -- a paradigm in which both the human and the robot share control [7, 8, 9, 10] -- has thus emerged as an ideal middle ground. By keeping the user in control of high-level decisions while delegating low-level actions to the autonomous robot, this approach both preserves user agency and enhances overall system performance. As such, assistive teleoperation is becoming a desirable paradigm for robots to serve as reliable partners in human-centric environments, such as assisting individuals with motor impairments [11, 12]. While promising, assistive teleoperation in everyday environments remains challenging. A longstanding challenge in assistive teleoperation is to infer human intents from user control inputs and assist users with correct actions [8]. This challenge is amplified in real-world settings, where robots must go beyond closed-set intent prediction [13, 14] to handle diverse, open-ended user goals across different contexts and scenes. As a result, a key capability the robot should possess is to interpret user control inputs within the visual context and infer intent through commonsense reasoning.
This teen 3D printed a beehive for his bedroom
Breakthroughs, discoveries, and DIY tips sent every weekday. While many 13-year-old boys might spend their summers playing video games or attending camp, Oliver Taylor decided to build a custom-made, 3D-printed beehive--in his bedroom. Oliver, who lives in Utah, built the DIY insect habitat with two hexagonal, 3D-printed units connected to his bedroom window. Bees enter through a ventilation tube attached to the window, which slightly resembles a stand-up air conditioning unit. The hexagonal hives are modular in design, meaning Oliver can theoretically continue expanding their size by connecting additional units.
The AI Birthday Letter That Blew Me Away
In May, I asked Google's chatbot, Gemini, to write a birthday letter to my best friend. Within seconds, it spat out the most impressive piece of AI writing I have ever encountered. Instead of reading as soulless, machine-generated text, the letter felt unnervingly like something I might've actually written. "You're probably rolling your eyes," the letter read, after a sentence that my friend would most definitely have rolled his eyes at. All I had typed into the chatbot was a nine-word prompt containing my friend's first name and the age he was turning.
Schools turn to handwritten exams as AI cheating surges
A growing number of fire departments across the country are turning to artificial intelligence to help detect and respond to wildfires more quickly. The rise of artificial intelligence in education is forcing schools and universities to rethink everything from homework policies to how final exams are administered. With tools like ChatGPT now widespread, students can generate essays, solve complex math problems or draft lab reports in seconds, raising urgent questions about what authentic learning looks like in 2025. To fight back, some schools are turning to an unlikely solution: pen and paper. The old-school "blue book," a lined booklet used for handwritten test answers, is staging a comeback, according to reporting from The Wall Street Journal.
Enhancing Clinical Multiple-Choice Questions Benchmarks with Knowledge Graph Guided Distractor Generation
Yang, Running, Deng, Wenlong, Chen, Minghui, Zhou, Yuyin, Li, Xiaoxiao
Clinical tasks such as diagnosis and treatment require strong decision-making abilities, highlighting the importance of rigorous evaluation benchmarks to assess the reliability of large language models (LLMs). In this work, we introduce a knowledge-guided data augmentation framework that enhances the difficulty of clinical multiple-choice question (MCQ) datasets by generating distractors (i.e., incorrect choices that are similar to the correct one and may confuse existing LLMs). Using our KG-based pipeline, the generated choices are both clinically plausible and deliberately misleading. Our approach involves multi-step, semantically informed walks on a medical knowledge graph to identify distractor paths-associations that are medically relevant but factually incorrect-which then guide the LLM in crafting more deceptive distractors. We apply the designed knowledge graph guided distractor generation (KGGDG) pipline, to six widely used medical QA benchmarks and show that it consistently reduces the accuracy of state-of-the-art LLMs. These findings establish KGGDG as a powerful tool for enabling more robust and diagnostic evaluations of medical LLMs.
DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment
Lu, Ke-Han, Chen, Zhehuai, Fu, Szu-Wei, Yang, Chao-Han Huck, Huang, Sung-Feng, Yang, Chih-Kai, Yu, Chee-En, Chen, Chun-Wei, Chen, Wei-Chih, Huang, Chien-yu, Lin, Yi-Cheng, Lin, Yu-Xiang, Fu, Chi-An, Kuan, Chun-Yi, Ren, Wenze, Chen, Xuanjun, Huang, Wei-Ping, Hu, En-Pei, Lin, Tzu-Quan, Wu, Yuan-Kuei, Huang, Kuan-Po, Huang, Hsiao-Ying, Chou, Huang-Cheng, Chang, Kai-Wei, Chiang, Cheng-Han, Ginsburg, Boris, Wang, Yu-Chiang Frank, Lee, Hung-yi
--We introduce DeST A2.5-Audio, a general-purpose Large Audio Language Model (LALM) designed for robust auditory perception and instruction-following, without requiring task-specific audio instruction-tuning. Recent LALMs typically augment Large Language Models (LLMs) with auditory capabilities by training on large-scale, manually curated or LLM-synthesized audio-instruction datasets. However, these approaches have often suffered from the catastrophic forgetting of the LLM's original language abilities. T o address this, we revisit the data construction pipeline and propose DeST A, a self-generated cross-modal alignment strategy in which the backbone LLM generates its own training targets. This approach preserves the LLM's native language proficiency while establishing effective audio-text alignment, thereby enabling zero-shot generalization without task-specific tuning. Using DeST A, we construct DeST A-AQA5M, a large-scale, task-agnostic dataset containing 5 million training samples derived from 7,000 hours of audio spanning 50 diverse datasets, including speech, environmental sounds, and music. DeST A2.5-Audio achieves state-of-the-art or competitive performance across a wide range of audio-language benchmarks, including Dynamic-SUPERB, MMAU, SAKURA, Speech-IFEval, and V oiceBench. Comprehensive comparative studies demonstrate that our self-generated strategy outperforms widely adopted data construction and training strategies in both auditory perception and instruction-following capabilities. Our findings underscore the importance of carefully designed data construction in LALM development and offer practical insights for building robust, general-purpose LALMs. HE development of general-purpose artificial intelligence has become a central focus in contemporary AI research, driven by the remarkable performance of large language models (LLMs) across various natural language understanding and generation tasks [1]-[7]. Building on these advancements, a promising direction is to equip LLMs with multi-modal understanding capabilities, leading to the emergence of Large Audio Language Models (LALMs) [8]-[22] and Large Vision Language Models (L VLMs) [23]-[27]. This paper focuses on building a general-purpose LALM, illustrated in Figure 1. To develop a general-purpose LALM, two core capabilities are essential: auditory perception and instruction-following. Auditory perception refers to the comprehensive processing of auditory information, including speech, non-verbal cues, background sounds, and music.
Do Role-Playing Agents Practice What They Preach? Belief-Behavior Consistency in LLM-Based Simulations of Human Trust
Mannekote, Amogh, Davies, Adam, Li, Guohao, Boyer, Kristy Elizabeth, Zhai, ChengXiang, Dorr, Bonnie J, Pinto, Francesco
As LLMs are increasingly studied as role-playing agents to generate synthetic data for human behavioral research, ensuring that their outputs remain coherent with their assigned roles has become a critical concern. In this paper, we investigate how consistently LLM-based role-playing agents' stated beliefs about the behavior of the people they are asked to role-play ("what they say") correspond to their actual behavior during role-play ("how they act"). Specifically, we establish an evaluation framework to rigorously measure how well beliefs obtained by prompting the model can predict simulation outcomes in advance. Using an augmented version of the GenAgents persona bank and the Trust Game (a standard economic game used to quantify players' trust and reciprocity), we introduce a belief-behavior consistency metric to systematically investigate how it is affected by factors such as: (1) the types of beliefs we elicit from LLMs, like expected outcomes of simulations versus task-relevant attributes of individual characters LLMs are asked to simulate; (2) when and how we present LLMs with relevant information about Trust Game; and (3) how far into the future we ask the model to forecast its actions. We also explore how feasible it is to impose a researcher's own theoretical priors in the event that the originally elicited beliefs are misaligned with research objectives. Our results reveal systematic inconsistencies between LLMs' stated (or imposed) beliefs and the outcomes of their role-playing simulation, at both an individual- and population-level. Specifically, we find that, even when models appear to encode plausible beliefs, they may fail to apply them in a consistent way. These findings highlight the need to identify how and when LLMs' stated beliefs align with their simulated behavior, allowing researchers to use LLM-based agents appropriately in behavioral studies.
Synthetic Heuristic Evaluation: A Comparison between AI- and Human-Powered Usability Evaluation
Zhong, Ruican, McDonald, David W., Hsieh, Gary
Usability evaluation is crucial in human-centered design but can be costly, requiring expert time and user compensation. In this work, we developed a method for synthetic heuristic evaluation using multimodal LLMs' ability to analyze images and provide design feedback. Comparing our synthetic evaluations to those by experienced UX practitioners across two apps, we found our evaluation identified 73% and 77% of usability issues, which exceeded the performance of 5 experienced human evaluators (57% and 63%). Compared to human evaluators, the synthetic evaluation's performance maintained consistent performance across tasks and excelled in detecting layout issues, highlighting potential attentional and perceptual strengths of synthetic evaluation. However, synthetic evaluation struggled with recognizing some UI components and design conventions, as well as identifying across screen violations. Additionally, testing synthetic evaluations over time and accounts revealed stable performance. Overall, our work highlights the performance differences between human and LLM-driven evaluations, informing the design of synthetic heuristic evaluations.