Agents
LLM Agents at the Roundtable: A Multi-Perspective and Dialectical Reasoning Framework for Essay Scoring
Jang, Jinhee, Moon, Ayoung, Jung, Minkyoung, Kim, YoungBin, Lee, Seung Jin
The emergence of large language models (LLMs) has brought a new paradigm to automated essay scoring (AES), a long-standing and practical application of natural language processing in education. However, achieving human-level multi-perspective understanding and judgment remains a challenge. In this work, we propose Roundtable Essay Scoring (RES), a multi-agent evaluation framework designed to perform precise and human-aligned scoring under a zero-shot setting. RES constructs evaluator agents based on LLMs, each tailored to a specific prompt and topic context. Each agent independently generates a trait-based rubric and conducts a multi-perspective evaluation. Then, by simulating a roundtable-style discussion, RES consolidates individual evaluations through a dialectical reasoning process to produce a final holistic score that more closely aligns with human evaluation. By enabling collaboration and consensus among agents with diverse evaluation perspectives, RES outperforms prior zero-shot AES approaches. Experiments on the ASAP dataset using ChatGPT and Claude show that RES achieves up to a 34.86% improvement in average QWK over straightforward prompting (Vanilla) methods.
Online Learning of Deceptive Policies under Intermittent Observation
Puthumanaillam, Gokul, Padmanabhan, Ram, Fuentes, Jose, Cruz, Nicole, Padrao, Paulo, Hernandez, Ruben, Jiang, Hao, Schafer, William, Bobadilla, Leonardo, Ornik, Melkior
In supervisory control settings, autonomous systems are not monitored continuously. Instead, monitoring often occurs at sporadic intervals within known bounds. We study the problem of deception, where an agent pursues a private objective while remaining plausibly compliant with a supervisor's reference policy when observations occur. Motivated by the behavior of real, human supervisors, we situate the problem within Theory of Mind: the representation of what an observer believes and expects to see. We show that Theory of Mind can be repurposed to steer online reinforcement learning (RL) toward such deceptive behavior. We model the supervisor's expectations and distill from them a single, calibrated scalar -- the expected evidence of deviation if an observation were to happen now. This scalar combines how unlike the reference and current action distributions appear, with the agent's belief that an observation is imminent. Injected as a state-dependent weight into a KL-regularized policy improvement step within an online RL loop, this scalar informs a closed-form update that smoothly trades off self-interest and compliance, thus sidestepping hand-crafted or heuristic policies. In real-world, real-time hardware experiments on marine (ASV) and aerial (UAV) navigation, our ToM-guided RL runs online, achieves high return and success with observed-trace evidence calibrated to the supervisor's expectations.
SWE-Effi: Re-Evaluating Software AI Agent System Effectiveness Under Resource Constraints
Fan, Zhiyu, Vasilevski, Kirill, Lin, Dayi, Chen, Boyuan, Chen, Yihao, Zhong, Zhiqing, Zhang, Jie M., He, Pinjia, Hassan, Ahmed E.
The advancement of large language models (LLMs) and code agents has demonstrated significant potential to assist software engineering (SWE) tasks, such as autonomous issue resolution and feature addition. Existing AI for software engineering leaderboards (e.g., SWE-bench) focus solely on solution accuracy, ignoring the crucial factor of effectiveness in a resource-constrained world. This is a universal problem that also exists beyond software engineering tasks: any AI system should be more than correct - it must also be cost-effective. To address this gap, we introduce SWE-Effi, a set of new metrics to re-evaluate AI systems in terms of holistic effectiveness scores. We define effectiveness as the balance between the accuracy of outcome (e.g., issue resolve rate) and the resources consumed (e.g., token and time). In this paper, we specifically focus on the software engineering scenario by re-ranking popular AI systems for issue resolution on a subset of the SWE-bench benchmark using our new multi-dimensional metrics. We found that AI system's effectiveness depends not just on the scaffold itself, but on how well it integrates with the base model, which is key to achieving strong performance in a resource-efficient manner. We also identified systematic challenges such as the "token snowball" effect and, more significantly, a pattern of "expensive failures". In these cases, agents consume excessive resources while stuck on unsolvable tasks - an issue that not only limits practical deployment but also drives up the cost of failed rollouts during RL training. Lastly, we observed a clear trade-off between effectiveness under the token budget and effectiveness under the time budget, which plays a crucial role in managing project budgets and enabling scalable reinforcement learning, where fast responses are essential.
LongCat-Flash Technical Report
Meituan LongCat Team, null, Bayan, null, Li, Bei, Lei, Bingye, Wang, Bo, Rong, Bolin, Wang, Chao, Zhang, Chao, Gao, Chen, Zhang, Chen, Sun, Cheng, Han, Chengcheng, Xi, Chenguang, Zhang, Chi, Peng, Chong, Qin, Chuan, Zhang, Chuyu, Chen, Cong, Wang, Congkui, Ma, Dan, Pan, Daoru, Bu, Defei, Zhao, Dengchang, Kong, Deyang, Liu, Dishan, Huo, Feiye, Li, Fengcun, Zhang, Fubao, Dong, Gan, Liu, Gang, Xu, Gang, Li, Ge, Tan, Guoqiang, Lin, Guoyuan, Jing, Haihang, Fu, Haomin, Yan, Haonan, Wen, Haoxing, Zhao, Haozhe, Liu, Hong, Shi, Hongmei, Hao, Hongyan, Tang, Hongyin, Lv, Huantian, Su, Hui, Li, Jiacheng, Liu, Jiahao, Li, Jiahuan, Yang, Jiajun, Wang, Jiaming, Yang, Jian, Tan, Jianchao, Sun, Jiaqi, Zhang, Jiaqi, Fu, Jiawei, Yang, Jiawei, Hu, Jiaxi, Qin, Jiayu, Wang, Jingang, He, Jiyuan, Kuang, Jun, Mei, Junhui, Liang, Kai, He, Ke, Zhang, Kefeng, Wang, Keheng, He, Keqing, Gao, Liang, Shi, Liang, Ma, Lianhui, Qiu, Lin, Kong, Lingbin, Si, Lingtong, Lyu, Linkun, Guo, Linsen, Yang, Liqi, Yan, Lizhi, Xia, Mai, Gao, Man, Zhang, Manyuan, Zhou, Meng, Shen, Mengxia, Tuo, Mingxiang, Zhu, Mingyang, Li, Peiguang, Pei, Peng, Zhao, Peng, Jia, Pengcheng, Sun, Pingwei, Gu, Qi, Li, Qianyun, Li, Qingyuan, Huang, Qiong, Duan, Qiyuan, Meng, Ran, Weng, Rongxiang, Shao, Ruichen, Li, Rumei, Wu, Shizhe, Liang, Shuai, Wang, Shuo, Dang, Suogui, Fang, Tao, Li, Tao, Chen, Tefeng, Bai, Tianhao, Zhou, Tianhao, Xie, Tingwen, He, Wei, Huang, Wei, Liu, Wei, Shi, Wei, Wang, Wei, Wu, Wei, Zhao, Weikang, Zan, Wen, Shi, Wenjie, Nan, Xi, Su, Xi, Li, Xiang, Mei, Xiang, Ji, Xiangyang, Xi, Xiangyu, Huang, Xiangzhou, Li, Xianpeng, Fu, Xiao, Liu, Xiao, Wei, Xiao, Cai, Xiaodong, Chen, Xiaolong, Liu, Xiaoqing, Li, Xiaotong, Shi, Xiaowei, Li, Xiaoyu, Wang, Xili, Chen, Xin, Hu, Xing, Miao, Xingyu, He, Xinyan, Zhang, Xuemiao, Hao, Xueyuan, Cao, Xuezhi, Cai, Xunliang, Yang, Xurui, Feng, Yan, Bai, Yang, Chen, Yang, Yang, Yang, Huo, Yaqi, Sun, Yerui, Lu, Yifan, Zhang, Yifan, Zang, Yipeng, Zhai, Yitao, Li, Yiyang, Yin, Yongjing, Lv, Yongkang, Zhou, Yongwei, Yang, Yu, Xie, Yuchen, Sun, Yueqing, Zheng, Yuewen, Wei, Yuhuai, Qian, Yulei, Liang, Yunfan, Tai, Yunfang, Zhao, Yunke, Yu, Zeyang, Zhang, Zhao, Yang, Zhaohua, Zhang, Zhenchao, Xia, Zhikang, Zou, Zhiye, Zeng, Zhizhao, Su, Zhongda, Chen, Zhuofan, Zhang, Zijian, Wang, Ziwen, Jiang, Zixu, Zhao, Zizhe, Wang, Zongyu, Su, Zunhai
We introduce LongCat-Flash, a 560-billion-parameter Mixture-of-Experts (MoE) language model designed for both computational efficiency and advanced agentic capabilities. Stemming from the need for scalable efficiency, LongCat-Flash adopts two novel designs: (a) Zero-computation Experts, which enables dynamic computational budget allocation and activates 18.6B-31.3B (27B on average) per token depending on contextual demands, optimizing resource usage. (b) Shortcut-connected MoE, which enlarges the computation-communication overlap window, demonstrating notable gains in inference efficiency and throughput compared to models of a comparable scale. We develop a comprehensive scaling framework for large models that combines hyperparameter transfer, model-growth initialization, a multi-pronged stability suite, and deterministic computation to achieve stable and reproducible training. Notably, leveraging the synergy among scalable architectural design and infrastructure efforts, we complete model training on more than 20 trillion tokens within 30 days, while achieving over 100 tokens per second (TPS) for inference at a cost of \$0.70 per million output tokens. To cultivate LongCat-Flash towards agentic intelligence, we conduct a large-scale pre-training on optimized mixtures, followed by targeted mid- and post-training on reasoning, code, and instructions, with further augmentation from synthetic data and tool use tasks. Comprehensive evaluations demonstrate that, as a non-thinking foundation model, LongCat-Flash delivers highly competitive performance among other leading models, with exceptional strengths in agentic tasks. The model checkpoint of LongCat-Flash is open-sourced to foster community research. LongCat Chat: https://longcat.ai Hugging Face: https://huggingface.co/meituan-longcat GitHub: https://github.com/meituan-longcat
The Anatomy of a Personal Health Agent
Heydari, A. Ali, Gu, Ken, Srinivas, Vidya, Yu, Hong, Zhang, Zhihan, Zhang, Yuwei, Paruchuri, Akshay, He, Qian, Palangi, Hamid, Hammerquist, Nova, Metwally, Ahmed A., Winslow, Brent, Kim, Yubin, Ayush, Kumar, Yang, Yuzhe, Narayanswamy, Girish, Xu, Maxwell A., Garrison, Jake, Lee, Amy Armento, Vafeiadou, Jenny, Graef, Ben, Galatzer-Levy, Isaac R., Schenck, Erik, Barakat, Andrew, Perez, Javier, Shreibati, Jacqueline, Hernandez, John, Faranesh, Anthony Z., Prieto, Javier L., Heneghan, Connor, Liu, Yun, Zhan, Jiening, Malhotra, Mark, Patel, Shwetak, Althoff, Tim, Liu, Xin, McDuff, Daniel, Xu, Xuhai "Orson"
Health is a fundamental pillar of human wellness, and the rapid advancements in large language models (LLMs) have driven the development of a new generation of health agents. However, the application of health agents to fulfill the diverse needs of individuals in daily non-clinical settings is underexplored. In this work, we aim to build a comprehensive personal health agent that is able to reason about multimodal data from everyday consumer wellness devices and common personal health records, and provide personalized health recommendations. To understand end-users' needs when interacting with such an assistant, we conducted an in-depth analysis of web search and health forum queries, alongside qualitative insights from users and health experts gathered through a user-centered design process. Based on these findings, we identified three major categories of consumer health needs, each of which is supported by a specialist sub-agent: (1) a data science agent that analyzes personal time-series wearable and health record data, (2) a health domain expert agent that integrates users' health and contextual data to generate accurate, personalized insights, and (3) a health coach agent that synthesizes data insights, guiding users using a specified psychological strategy and tracking users' progress. Furthermore, we propose and develop the Personal Health Agent (PHA), a multi-agent framework that enables dynamic, personalized interactions to address individual health needs. To evaluate each sub-agent and the multi-agent system, we conducted automated and human evaluations across 10 benchmark tasks, involving more than 7,000 annotations and 1,100 hours of effort from health experts and end-users. Our work represents the most comprehensive evaluation of a health agent to date and establishes a strong foundation towards the futuristic vision of a personal health agent accessible to everyone.
Disproving the Feasibility of Learned Confidence Calibration Under Binary Supervision: An Information-Theoretic Impossibility
Nair, Arjun S., Sinaga, Kristina P.
We prove a fundamental impossibility theorem: neural networks cannot simultaneously learn well-calibrated confidence estimates with meaningful diversity when trained using binary correct/incorrect supervision. Through rigorous mathematical analysis and comprehensive empirical evaluation spanning negative reward training, symmetric loss functions, and post-hoc calibration methods, we demonstrate this is an information-theoretic constraint, not a methodological failure. Our experiments reveal universal failure patterns: negative rewards produce extreme underconfidence (ECE greater than 0.8) while destroying confidence diversity (std less than 0.05), symmetric losses fail to escape binary signal averaging, and post-hoc methods achieve calibration (ECE less than 0.02) only by compressing the confidence distribution. We formalize this as an underspecified mapping problem where binary signals cannot distinguish between different confidence levels for correct predictions: a 60 percent confident correct answer receives identical supervision to a 90 percent confident one. Crucially, our real-world validation shows 100 percent failure rate for all training methods across MNIST, Fashion-MNIST, and CIFAR-10, while post-hoc calibration's 33 percent success rate paradoxically confirms our theorem by achieving calibration through transformation rather than learning. This impossibility directly explains neural network hallucinations and establishes why post-hoc calibration is mathematically necessary, not merely convenient. We propose novel supervision paradigms using ensemble disagreement and adaptive multi-agent learning that could overcome these fundamental limitations without requiring human confidence annotations.
The Sum Leaks More Than Its Parts: Compositional Privacy Risks and Mitigations in Multi-Agent Collaboration
Patil, Vaidehi, Stengel-Eskin, Elias, Bansal, Mohit
As large language models (LLMs) become integral to multi-agent systems, new privacy risks emerge that extend beyond memorization, direct inference, or single-turn evaluations. In particular, seemingly innocuous responses, when composed across interactions, can cumulatively enable adversaries to recover sensitive information, a phenomenon we term compositional privacy leakage. We present the first systematic study of such compositional privacy leaks and possible mitigation methods in multi-agent LLM systems. First, we develop a framework that models how auxiliary knowledge and agent interactions jointly amplify privacy risks, even when each response is benign in isolation. Next, to mitigate this, we propose and evaluate two defense strategies: (1) Theory-of-Mind defense (ToM), where defender agents infer a questioner's intent by anticipating how their outputs may be exploited by adversaries, and (2) Collaborative Consensus Defense (CoDef), where responder agents collaborate with peers who vote based on a shared aggregated state to restrict sensitive information spread. Crucially, we balance our evaluation across compositions that expose sensitive information and compositions that yield benign inferences. Our experiments quantify how these defense strategies differ in balancing the privacy-utility trade-off. We find that while chain-of-thought alone offers limited protection to leakage (~39% sensitive blocking rate), our ToM defense substantially improves sensitive query blocking (up to 97%) but can reduce benign task success. CoDef achieves the best balance, yielding the highest Balanced Outcome (79.8%), highlighting the benefit of combining explicit reasoning with defender collaboration. Together, our results expose a new class of risks in collaborative LLM deployments and provide actionable insights for designing safeguards against compositional, context-driven privacy leakage.
Continuous-Time Value Iteration for Multi-Agent Reinforcement Learning
Wang, Xuefeng, Zhang, Lei, Pu, Henglin, Qureshi, Ahmed H., Li, Husheng
Existing reinforcement learning (RL) methods struggle with complex dynamical systems that demand interactions at high frequencies or irregular time intervals. Continuous-time RL (CTRL) has emerged as a promising alternative by replacing discrete-time Bellman recursion with differential value functions defined as viscosity solutions of the Hamilton--Jacobi--Bellman (HJB) equation. While CTRL has shown promise, its applications have been largely limited to the single-agent domain. This limitation stems from two key challenges: (i) conventional solution methods for HJB equations suffer from the curse of dimensionality (CoD), making them intractable in high-dimensional systems; and (ii) even with HJB-based learning approaches, accurately approximating centralized value functions in multi-agent settings remains difficult, which in turn destabilizes policy training. In this paper, we propose a CT-MARL framework that uses physics-informed neural networks (PINNs) to approximate HJB-based value functions at scale. To ensure the value is consistent with its differential structure, we align value learning with value-gradient learning by introducing a Value Gradient Iteration (VGI) module that iteratively refines value gradients along trajectories. This improves gradient fidelity, in turn yielding more accurate values and stronger policy learning. We evaluate our method using continuous-time variants of standard benchmarks, including multi-agent particle environment (MPE) and multi-agent MuJoCo. Our results demonstrate that our approach consistently outperforms existing continuous-time RL baselines and scales to complex multi-agent dynamics.
Judging with Many Minds: Do More Perspectives Mean Less Prejudice? On Bias Amplifications and Resistance in Multi-Agent Based LLM-as-Judge
Ma, Chiyu, Zhang, Enpei, Zhao, Yilun, Liu, Wenjun, Jia, Yaning, Qing, Peijun, Shi, Lin, Cohan, Arman, Yan, Yujun, Vosoughi, Soroush
LLM-as-Judge has emerged as a scalable alternative to human evaluation, enabling large language models (LLMs) to provide reward signals in trainings. While recent work has explored multi-agent extensions such as multi-agent debate and meta-judging to enhance evaluation quality, the question of how intrinsic biases manifest in these settings remains underexplored. In this study, we conduct a systematic analysis of four diverse bias types: position bias, verbosity bias, chain-of-thought bias, and bandwagon bias. We evaluate these biases across two widely adopted multi-agent LLM-as-Judge frameworks: Multi-Agent-Debate and LLM-as-Meta-Judge. Our results show that debate framework amplifies biases sharply after the initial debate, and this increased bias is sustained in subsequent rounds, while meta-judge approaches exhibit greater resistance. We further investigate the incorporation of PINE, a leading single-agent debiasing method, as a bias-free agent within these systems. The results reveal that this bias-free agent effectively reduces biases in debate settings but provides less benefit in meta-judge scenarios. Our work provides a comprehensive study of bias behavior in multi-agent LLM-as-Judge systems and highlights the need for targeted bias mitigation strategies in collaborative evaluation settings.
Trustless Autonomy: Understanding Motivations, Benefits, and Governance Dilemmas in Self-Sovereign Decentralized AI Agents
Hu, Botao Amber, Liu, Yuhan, Rong, Helena
The recent trend of self-sovereign Decentralized AI Agents (DeAgents) combines Large Language Model (LLM)-based AI agents with decentralization technologies such as blockchain smart contracts and trusted execution environments (TEEs). These tamper-resistant trustless substrates allow agents to achieve self-sovereignty through ownership of cryptowallet private keys and control of digital assets and social media accounts. DeAgents eliminate centralized control and reduce human intervention, addressing key trust concerns inherent in centralized AI systems. This contributes to social computing by enabling new human cooperative paradigm "intelligence as commons." However, given ongoing challenges in LLM reliability such as hallucinations, this creates paradoxical tension between trustlessness and unreliable autonomy. This study addresses this empirical research gap through interviews with DeAgents stakeholders-experts, founders, and developers-to examine their motivations, benefits, and governance dilemmas. The findings will guide future DeAgents system and protocol design and inform discussions about governance in sociotechnical AI systems in the future agentic web.