Large Language Model
What is an "Abstract Reasoner"? Revisiting Experiments and Arguments about Large Language Models
Yun, Tian, Sun, Chen, Pavlick, Ellie
Recent work has argued that large language models (LLMs) are not "abstract reasoners", citing their poor zero-shot performance on a variety of challenging tasks as evidence. We revisit these experiments in order to add nuance to the claim. First, we show that while LLMs indeed perform poorly in a zero-shot setting, even tuning a small subset of parameters for input encoding can enable near-perfect performance. However, we also show that this finetuning does not necessarily transfer across datasets. We take this collection of empirical results as an invitation to (re-)open the discussion of what it means to be an "abstract reasoner", and why it matters whether LLMs fit the bill.
Systematic Evaluation of Knowledge Graph Repair with Large Language Models
Lin, Tung-Wei, Fierro, Gabe, Li, Han, Hong, Tianzhen, Nuzzo, Pierluigi, Sangiovanni-Vinentelli, Alberto
We present a systematic approach for evaluating the quality of knowledge graph repairs with respect to constraint violations defined in shapes constraint language (SHACL). Current evaluation methods rely on \emph{ad hoc} datasets, which limits the rigorous analysis of repair systems in more general settings. Our method addresses this gap by systematically generating violations using a novel mechanism, termed violation-inducing operations (VIOs). We use the proposed evaluation framework to assess a range of repair systems which we build using large language models. We analyze the performance of these systems across different prompting strategies. Results indicate that concise prompts containing both the relevant violated SHACL constraints and key contextual information from the knowledge graph yield the best performance.
From Articles to Code: On-Demand Generation of Core Algorithms from Scientific Publications
Movassaghi, Cameron S., Momenzadeh, Amanda, Meyer, Jesse G.
Maintaining software packages imposes significant costs due to dependency management, bug fixes, and versioning. We show that rich method descriptions in scientific publications can serve as standalone specifications for modern large language models (LLMs), enabling on-demand code generation that could supplant human-maintained libraries. We benchmark state-of-the-art models (GPT-o4-mini-high, Gemini Pro 2.5, Claude Sonnet 4) by tasking them with implementing a diverse set of core algorithms drawn from original publications. Our results demonstrate that current LLMs can reliably reproduce package functionality with performance indistinguishable from conventional libraries. These findings foreshadow a paradigm shift toward flexible, on-demand code generation and away from static, human-maintained packages, which will result in reduced maintenance overhead by leveraging published articles as sufficient context for the automated implementation of analytical workflows.
Intent Recognition and Out-of-Scope Detection using LLMs in Multi-party Conversations
Castillo-Lรณpez, Galo, de Chalendar, Gaรซl, Semmar, Nasredine
Intent recognition is a fundamental component in task-oriented dialogue systems (TODS). Determining user intents and detecting whether an intent is Out-of-Scope (OOS) is crucial for TODS to provide reliable responses. However, traditional TODS require large amount of annotated data. In this work we propose a hybrid approach to combine BERT and LLMs in zero and few-shot settings to recognize intents and detect OOS utterances. Our approach leverages LLMs generalization power and BERT's computational efficiency in such scenarios. We evaluate our method on multi-party conversation corpora and observe that sharing information from BERT outputs to LLMs leads to system performance improvement.
CoEx -- Co-evolving World-model and Exploration
Planning in modern LLM agents relies on the utilization of LLM as an internal world model, acquired during pretraining. However, existing agent designs fail to effectively assimilate new observations into dynamic updates of the world model. This reliance on the LLM's static internal world model is progressively prone to misalignment with the underlying true state of the world, leading to the generation of divergent and erroneous plans. We introduce a hierarchical agent architecture, CoEx, in which hierarchical state abstraction allows LLM planning to co-evolve with a dynamically updated model of the world. CoEx plans and interacts with the world by using LLM reasoning to orchestrate dynamic plans consisting of subgoals, and its learning mechanism continuously incorporates these subgoal experiences into a persistent world model in the form of a neurosymbolic belief state, comprising textual inferences and code-based symbolic memory. We evaluate our agent across a diverse set of agent scenarios involving rich environments and complex tasks including ALFWorld, PDDL, and Jericho. Our experiments show that CoEx outperforms existing agent paradigms in planning and exploration.
Promoting Online Safety by Simulating Unsafe Conversations with LLMs
Hoffman, Owen, Peng, Kangze, You, Zehua, Kamal, Sajid, Venkatagiri, Sukrit
Generative AI, including large language models (LLMs) have the potential -- and already are being used -- to increase the speed, scale, and types of unsafe conversations online. LLMs lower the barrier for entry for bad actors to create unsafe conversations in particular because of their ability to generate persuasive and human-like text. In our current work, we explore ways to promote online safety by teaching people about unsafe conversations that can occur online with and without LLMs. We build on prior work that shows that LLMs can successfully simulate scam conversations. We also leverage research in the learning sciences that shows that providing feedback on one's hypothetical actions can promote learning. In particular, we focus on simulating scam conversations using LLMs. Our work incorporates two LLMs that converse with each other to simulate realistic, unsafe conversations that people may encounter online between a scammer LLM and a target LLM but users of our system are asked provide feedback to the target LLM.
SmartCLIP: Modular Vision-language Alignment with Identification Guarantees
Xie, Shaoan, Kong, Lingjing, Zheng, Yujia, Yao, Yu, Tang, Zeyu, Xing, Eric P., Chen, Guangyi, Zhang, Kun
Contrastive Language-Image Pre-training (CLIP) [37] has emerged as a pivotal model in computer vision and multi-modal learning, achieving state-of-the-art performance at aligning visual and textual representations through contrastive learning. However, CLIP struggles with potential information misalignment in many image-text datasets and suffers from entangled representation. On the one hand, short captions for a single image in datasets like MSCOCO may describe disjoint regions in the image, leaving the model uncertain about which visual features to retain or disregard. On the other hand, directly aligning long captions with images can lead to the retention of entangled details, preventing the model from learning disentangled, atomic concepts - ultimately limiting its generalization on certain downstream tasks involving short prompts. In this paper, we establish theoretical conditions that enable flexible alignment between textual and visual representations across varying levels of granularity. Specifically, our framework ensures that a model can not only preserve cross-modal semantic information in its entirety but also disentangle visual representations to capture fine-grained textual concepts. Building on this foundation, we introduce SmartCLIP, a novel approach that identifies and aligns the most relevant visual and textual representations in a modular manner . Superior performance across various tasks demonstrates its capability to handle information misalignment and supports our identification theory.
ChatGPT gets 'study mode' to guide students without spoon-feeding answers
OpenAI has launched a new "study mode" for ChatGPT that's designed to help students better understand complex topics--but instead of dishing out direct answers, study mode employs the Socratic method to ask questions and guide users to finding those answers. Or another way to look at it: in study mode conversations, ChatGPT gradually rolls out information to the user in stages to avoid overloading and overwhelming, and to prevent the AI chatbot from doing all the work on the user's behalf. According to OpenAI, study mode was developed in collaboration with teachers, researchers, and education experts. It's based on customized system instructions rather than an entirely new AI model. Study mode will first be available to users on ChatGPT Free, Plus, Pro, and Team plans. ChatGPT Edu users will get access within a few weeks.
The Download: a 30-year old baby, and OpenAI's push into colleges
A baby boy has just won the new record for the "oldest baby." Thaddeus Daniel Pierce, who arrived on July 26, developed from an embryo that had been in storage for 30 and a half years. Lindsey and her husband, Tim Pierce, who live in London, Ohio, "adopted" the embryo from Linda Archerd, who had it created in 1994. The couple, aged 35 and 34, respectively, had been trying for a baby for seven years. OpenAI is launching Study Mode, a version of ChatGPT for college students that it promises will act less like a lookup tool and more like a friendly, always-available tutor.
9 creative ways to use ChatGPT that are outside the box
ChatGPT is quite capable at considering both sides of an argument in a logical way, which makes it the perfect tool when you need a third party to mediate an argument--whether that argument is between you and someone else, or just one you're debating in your own head. As an example, I asked ChatGPT to help me look at both sides of the debate around investing more money in space travel. Some people see humanity reaching out to the stars as the best way forward, while others see it as a waste of money. ChatGPT looked at the pros and cons of each side and ended with a sensible set of conclusions. However, it did end up sitting on the fence a bit, so you may need to push it harder if you want a strong answer one way or the other.