Large Language Model
SWE-Bench-CL: Continual Learning for Coding Agents
Joshi, Thomas, Chowdhury, Shayan, Uysal, Fatih
Large Language Models (LLMs) have achieved impressive results on static code-generation benchmarks, but real-world software development unfolds as a continuous stream of evolving issues, fixes, and feature requests. We introduce SWE-Bench-CL, a novel continual learning benchmark built on the human-verified SWE-Bench Verified dataset introduced by OpenAI and Princeton-NLP in 2024. By organizing GitHub issues into chronologically ordered sequences that reflect natural repository evolution, SWE-Bench-CL enables direct evaluation of an agent's ability to accumulate experience, transfer knowledge across tasks, and resist catastrophic forgetting. We complement the dataset with (i) a preliminary analysis of inter-task structural similarity and contextual sensitivity, (ii) an interactive LangGraph-based evaluation framework augmented with a FAISS-backed semantic memory module, and (iii) a suite of specialized continual learning metrics -- including average accuracy, forgetting, forward/backward transfer, tool-use efficiency, and a generalized Composite Continual Learning Score and CL-F-beta score -- to capture the stability-plasticity trade-off. We outline a rigorous experimental protocol comparing memory-enabled and memory-disabled agents across diverse Python repositories. All code and data are publicly available at https://github.com/thomasjoshi/agents-never-forget, providing the community with a reproducible platform for developing more adaptive and robust AI agents in software engineering.
Integrating Universal Generative AI Platforms in Educational Labs to Foster Critical Thinking and Digital Literacy
Znamenskiy, Vasiliy, Niyazov, Rafael, Hernandez, Joel
This paper presents a new educational framework for integrating generative artificial intelligence (GenAI) platforms such as ChatGPT, Claude, and Gemini into laboratory activities aimed at developing critical thinking and digital literacy among undergraduate students. Recognizing the limitations and risks of uncritical reliance on large language models (LLMs), the proposed pedagogical model reframes GenAI as a research subject and cognitive tool. Students formulate discipline-specific prompts and evaluate GenAI-generated responses in text, image, and video modalities. A pilot implementation in a general astronomy course for non-science majors demonstrated high levels of engagement and critical reflection, with many students continuing the activity after class and presenting results at a research symposium. The results highlight the importance of structured AI interactions in education and suggest that GenAI can improve learning outcomes when combined with reflective assessment methods. The study proposes a replicable model for interdisciplinary AI-integrated lab work, adaptable to scientific disciplines. See the guide to learning activities based on Generative-Ai platforms: https://doi.org/10.5281/zenodo.15555802
Intertextual Parallel Detection in Biblical Hebrew: A Transformer-Based Benchmark
Identifying parallel passages in biblical Hebrew (BH) is central to biblical scholarship for understanding intertextual relationships. Traditional methods rely on manual comparison, a labor-intensive process prone to human error. This study evaluates the potential of pre-trained transformer-based language models, including E5, AlephBERT, MPNet, and LaBSE, for detecting textual parallels in the Hebrew Bible. Focusing on known parallels between Samuel/Kings and Chronicles, I assessed each model's capability to generate word embeddings distinguishing parallel from non-parallel passages. Using cosine similarity and Wasserstein Distance measures, I found that E5 and AlephBERT show promise; E5 excels in parallel detection, while AlephBERT demonstrates stronger non-parallel differentiation. These findings indicate that pre-trained models can enhance the efficiency and accuracy of detecting intertextual parallels in ancient texts, suggesting broader applications for ancient language studies.
Ovis-U1 Technical Report
Wang, Guo-Hua, Zhao, Shanshan, Zhang, Xinjie, Cao, Liangfu, Zhan, Pengxin, Duan, Lunhao, Lu, Shiyin, Fu, Minghao, Chen, Xiaohao, Zhao, Jianshan, Li, Yang, Chen, Qing-Guo
In this report, we introduce Ovis-U1, a 3-billion-parameter unified model that integrates multimodal understanding, text-to-image generation, and image editing capabilities. Building on the foundation of the Ovis series, Ovis-U1 incorporates a diffusion-based visual decoder paired with a bidirectional token refiner, enabling image generation tasks comparable to leading models like GPT-4o. Unlike some previous models that use a frozen MLLM for generation tasks, Ovis-U1 utilizes a new unified training approach starting from a language model. Compared to training solely on understanding or generation tasks, unified training yields better performance, demonstrating the enhancement achieved by integrating these two tasks. Ovis-U1 achieves a score of 69.6 on the OpenCompass Multi-modal Academic Benchmark, surpassing recent state-of-the-art models such as Ristretto-3B and SAIL-VL-1.5-2B. In text-to-image generation, it excels with scores of 83.72 and 0.89 on the DPG-Bench and GenEval benchmarks, respectively. For image editing, it achieves 4.00 and 6.42 on the ImgEdit-Bench and GEdit-Bench-EN, respectively. As the initial version of the Ovis unified model series, Ovis-U1 pushes the boundaries of multimodal understanding, generation, and editing.
Text Production and Comprehension by Human and Artificial Intelligence: Interdisciplinary Workshop Report
This report synthesizes the outcomes of a recent interdisciplinary workshop that brought together leading experts in cognitive psychology, language learning, and artificial intelligence (AI)-based natural language processing (NLP). The workshop, funded by the National Science Foundation, aimed to address a critical knowledge gap in our understanding of the relationship between AI language models and human cognitive processes in text comprehension and composition. Through collaborative dialogue across cognitive, linguistic, and technological perspectives, workshop participants examined the underlying processes involved when humans produce and comprehend text, and how AI can both inform our understanding of these processes and augment human capabilities. The workshop revealed emerging patterns in the relationship between large language models (LLMs) and human cognition, with highlights on both the capabilities of LLMs and their limitations in fully replicating human-like language understanding and generation. Key findings include the potential of LLMs to offer insights into human language processing, the increasing alignment between LLM behavior and human language processing when models are fine-tuned with human feedback, and the opportunities and challenges presented by human-AI collaboration in language tasks. By synthesizing these findings, this report aims to guide future research, development, and implementation of LLMs in cognitive psychology, linguistics, and education. It emphasizes the importance of ethical considerations and responsible use of AI technologies while striving to enhance human capabilities in text comprehension and production through effective human-AI collaboration.
Seamless Interaction: Dyadic Audiovisual Motion Modeling and Large-Scale Dataset
Agrawal, Vasu, Akinyemi, Akinniyi, Alvero, Kathryn, Behrooz, Morteza, Buffalini, Julia, Carlucci, Fabio Maria, Chen, Joy, Chen, Junming, Chen, Zhang, Cheng, Shiyang, Chowdary, Praveen, Chuang, Joe, D'Avirro, Antony, Daly, Jon, Dong, Ning, Duppenthaler, Mark, Gao, Cynthia, Girard, Jeff, Gleize, Martin, Gomez, Sahir, Gong, Hongyu, Govindarajan, Srivathsan, Han, Brandon, He, Sen, Hernandez, Denise, Hristov, Yordan, Huang, Rongjie, Inaguma, Hirofumi, Jain, Somya, Janardhan, Raj, Jia, Qingyao, Klaiber, Christopher, Kovachev, Dejan, Kumar, Moneish, Li, Hang, Li, Yilei, Litvin, Pavel, Liu, Wei, Ma, Guangyao, Ma, Jing, Ma, Martin, Ma, Xutai, Mantovani, Lucas, Miglani, Sagar, Mohan, Sreyas, Morency, Louis-Philippe, Ng, Evonne, Ng, Kam-Woh, Nguyen, Tu Anh, Oberai, Amia, Peloquin, Benjamin, Pino, Juan, Popovic, Jovan, Poursaeed, Omid, Prada, Fabian, Rakotoarison, Alice, Ranjan, Rakesh, Richard, Alexander, Ropers, Christophe, Saleem, Safiyyah, Sharma, Vasu, Shcherbyna, Alex, Shen, Jia, Shen, Jie, Stathopoulos, Anastasis, Sun, Anna, Tomasello, Paden, Tran, Tuan, Turkatenko, Arina, Wan, Bo, Wang, Chao, Wang, Jeff, Williamson, Mary, Wood, Carleigh, Xiang, Tao, Yang, Yilin, Yao, Julien, Zhang, Chen, Zhang, Jiemin, Zhang, Xinyue, Zheng, Jason, Zhyzheria, Pavlo, Zikes, Jan, Zollhoefer, Michael
Human communication involves a complex interplay of verbal and nonverbal signals, essential for conveying meaning and achieving interpersonal goals. To develop socially intelligent AI technologies, it is crucial to develop models that can both comprehend and generate dyadic behavioral dynamics. To this end, we introduce the Seamless Interaction Dataset, a large-scale collection of over 4,000 hours of face-to-face interaction footage from over 4,000 participants in diverse contexts. This dataset enables the development of AI technologies that understand dyadic embodied dynamics, unlocking breakthroughs in virtual agents, telepresence experiences, and multimodal content analysis tools. We also develop a suite of models that utilize the dataset to generate dyadic motion gestures and facial expressions aligned with human speech. These models can take as input both the speech and visual behavior of their interlocutors. We present a variant with speech from an LLM model and integrations with 2D and 3D rendering methods, bringing us closer to interactive virtual agents. Additionally, we describe controllable variants of our motion models that can adapt emotional responses and expressivity levels, as well as generating more semantically-relevant gestures. Finally, we discuss methods for assessing the quality of these dyadic motion models, which are demonstrating the potential for more intuitive and responsive human-AI interactions.
The Singapore Consensus on Global AI Safety Research Priorities
Bengio, Yoshua, Maharaj, Tegan, Ong, Luke, Russell, Stuart, Song, Dawn, Tegmark, Max, Xue, Lan, Zhang, Ya-Qin, Casper, Stephen, Lee, Wan Sie, Mindermann, Sรถren, Wilfred, Vanessa, Balachandran, Vidhisha, Barez, Fazl, Belinsky, Michael, Bello, Imane, Bourgon, Malo, Brakel, Mark, Campos, Simรฉon, Cass-Beggs, Duncan, Chen, Jiahao, Chowdhury, Rumman, Seah, Kuan Chua, Clune, Jeff, Dai, Juntao, Delaborde, Agnes, Dziri, Nouha, Eiras, Francisco, Engels, Joshua, Fan, Jinyu, Gleave, Adam, Goodman, Noah, Heide, Fynn, Heidecke, Johannes, Hendrycks, Dan, Hodes, Cyrus, Hsiang, Bryan Low Kian, Huang, Minlie, Jawhar, Sami, Jingyu, Wang, Kalai, Adam Tauman, Kamphuis, Meindert, Kankanhalli, Mohan, Kantamneni, Subhash, Kirk, Mathias Bonde, Kwa, Thomas, Ladish, Jeffrey, Lam, Kwok-Yan, Sie, Wan Lee, Lee, Taewhi, Li, Xiaojian, Liu, Jiajun, Lu, Chaochao, Mai, Yifan, Mallah, Richard, Michael, Julian, Moรซs, Nick, Mรถller, Simon, Nam, Kihyuk, Ng, Kwan Yee, Nitzberg, Mark, Nushi, Besmira, hรigeartaigh, Seรกn O, Ortega, Alejandro, Peignรฉ, Pierre, Petrie, James, Prud'Homme, Benjamin, Rabbany, Reihaneh, Sanchez-Pi, Nayat, Schwettmann, Sarah, Shlegeris, Buck, Siddiqui, Saad, Sinha, Aradhana, Soto, Martรญn, Tan, Cheston, Ting, Dong, Tjhi, William, Trager, Robert, Tse, Brian, H., Anthony Tung K., Wilfred, Vanessa, Willes, John, Wong, Denise, Xu, Wei, Xu, Rongwu, Zeng, Yi, Zhang, HongJiang, ลฝikeliฤ, Djordje
Rapidly improving AI capabilities and autonomy hold significant promise of transformation, but are also driving vigorous debate on how to ensure that AI is safe, i.e., trustworthy, reliable, and secure. Building a trusted ecosystem is therefore essential -- it helps people embrace AI with confidence and gives maximal space for innovation while avoiding backlash. The "2025 Singapore Conference on AI (SCAI): International Scientific Exchange on AI Safety" aimed to support research in this space by bringing together AI scientists across geographies to identify and synthesise research priorities in AI safety. This resulting report builds on the International AI Safety Report chaired by Yoshua Bengio and backed by 33 governments. By adopting a defence-in-depth model, this report organises AI safety research domains into three types: challenges with creating trustworthy AI systems (Development), challenges with evaluating their risks (Assessment), and challenges with monitoring and intervening after deployment (Control).
Benchmarking the Pedagogical Knowledge of Large Language Models
Leliรจvre, Maxime, Waldock, Amy, Liu, Meng, Aspillaga, Natalia Valdรฉs, Mackintosh, Alasdair, Portela, Marรญa Josรฉ Ogando, Lee, Jared, Atherton, Paul, Ince, Robin A. A., Garrod, Oliver G. B.
Benchmarks like Massive Multitask Language Understanding (MMLU) have played a pivotal role in evaluating AI's knowledge and abilities across diverse domains. However, existing benchmarks predominantly focus on content knowledge, leaving a critical gap in assessing models' understanding of pedagogy - the method and practice of teaching. This paper introduces The Pedagogy Benchmark, a novel dataset designed to evaluate large language models on their Cross-Domain Pedagogical Knowledge (CDPK) and Special Education Needs and Disability (SEND) pedagogical knowledge. These benchmarks are built on a carefully curated set of questions sourced from professional development exams for teachers, which cover a range of pedagogical subdomains such as teaching strategies and assessment methods. Here we outline the methodology and development of these benchmarks. We report results for 97 models, with accuracies spanning a range from 28% to 89% on the pedagogical knowledge questions. We consider the relationship between cost and accuracy and chart the progression of the Pareto value frontier over time. We provide online leaderboards at https://rebrand.ly/pedagogy which are updated with new models and allow interactive exploration and filtering based on various model properties, such as cost per token and open-vs-closed weights, as well as looking at performance in different subjects. LLMs and generative AI have tremendous potential to influence education and help to address the global learning crisis. Education-focused benchmarks are crucial to measure models' capacities to understand pedagogical concepts, respond appropriately to learners' needs, and support effective teaching practices across diverse contexts. They are needed for informing the responsible and evidence-based deployment of LLMs and LLM-based tools in educational settings, and for guiding both development and policy decisions.
Not Minds, but Signs: Reframing LLMs through Semiotics
This paper challenges the prevailing tendency to frame Large Language Models (LLMs) as cognitive systems, arguing instead for a semiotic perspective that situates these models within the broader dynamics of sign manipulation and meaning-making. Rather than assuming that LLMs understand language or simulate human thought, we propose that their primary function is to recombine, recontextualize, and circulate linguistic forms based on probabilistic associations. By shifting from a cognitivist to a semiotic framework, we avoid anthropomorphism and gain a more precise understanding of how LLMs participate in cultural processes, not by thinking, but by generating texts that invite interpretation. Through theoretical analysis and practical examples, the paper demonstrates how LLMs function as semiotic agents whose outputs can be treated as interpretive acts, open to contextual negotiation and critical reflection. We explore applications in literature, philosophy, education, and cultural production, emphasizing how LLMs can serve as tools for creativity, dialogue, and critical inquiry. The semiotic paradigm foregrounds the situated, contingent, and socially embedded nature of meaning, offering a more rigorous and ethically aware framework for studying and using LLMs. Ultimately, this approach reframes LLMs as technological participants in an ongoing ecology of signs. They do not possess minds, but they alter how we read, write, and make meaning, compelling us to reconsider the foundations of language, interpretation, and the role of artificial systems in the production of knowledge.
Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples
Recent advancements in audio-aware large language models (ALLMs) enable them to process and understand audio inputs. However, these models often hallucinate non-existent sound events, reducing their reliability in real-world applications. To address this, we propose LISTEN (Learning to Identify Sounds Through Extended Negative Samples), a contrastive-like training method that enhances ALLMs' ability to distinguish between present and absent sounds using synthesized data from the backbone LLM. Unlike prior approaches, our method requires no modification to LLM parameters and efficiently integrates audio representations via a lightweight adapter. Experiments show that LISTEN effectively mitigates hallucinations while maintaining impressive performance on existing audio question and reasoning benchmarks. At the same time, it is more efficient in both data and computation.