Personal
Interactive Debugging and Steering of Multi-Agent AI Systems
Epperson, Will, Bansal, Gagan, Dibia, Victor, Fourney, Adam, Gerrits, Jack, Zhu, Erkang, Amershi, Saleema
Fully autonomous teams of LLM-powered AI agents are emerging that collaborate to perform complex tasks for users. What challenges do developers face when trying to build and debug these AI agent teams? In formative interviews with five AI agent developers, we identify core challenges: difficulty reviewing long agent conversations to localize errors, lack of support in current tools for interactive debugging, and the need for tool support to iterate on agent configuration. Based on these needs, we developed an interactive multi-agent debugging tool, AGDebugger, with a UI for browsing and sending messages, the ability to edit and reset prior agent messages, and an overview visualization for navigating complex message histories. In a two-part user study with 14 participants, we identify common user strategies for steering agents and highlight the importance of interactive message resets for debugging. Our studies deepen understanding of interfaces for debugging increasingly important agentic workflows.
Can AI Model the Complexities of Human Moral Decision-Making? A Qualitative Study of Kidney Allocation Decisions
Keswani, Vijay, Conitzer, Vincent, Sinnott-Armstrong, Walter, Nguyen, Breanna K., Heidari, Hoda, Borg, Jana Schaich
A growing body of work in Ethical AI attempts to capture human moral judgments through simple computational models. The key question we address in this work is whether such simple AI models capture {the critical} nuances of moral decision-making by focusing on the use case of kidney allocation. We conducted twenty interviews where participants explained their rationale for their judgments about who should receive a kidney. We observe participants: (a) value patients' morally-relevant attributes to different degrees; (b) use diverse decision-making processes, citing heuristics to reduce decision complexity; (c) can change their opinions; (d) sometimes lack confidence in their decisions (e.g., due to incomplete information); and (e) express enthusiasm and concern regarding AI assisting humans in kidney allocation decisions. Based on these findings, we discuss challenges of computationally modeling moral judgments {as a stand-in for human input}, highlight drawbacks of current approaches, and suggest future directions to address these issues.
Congratulations to the #AAAI2025 outstanding paper award winners
The AAAI 2025 outstanding paper awards were announced during the opening ceremony of the 39th Annual AAAI Conference on Artificial Intelligence on Thursday 27 February. Papers are recommended for consideration during the review process by members of the Program Committee. This year, three papers have been selected as outstanding papers, with a further paper being recognised in the special track on AI for social impact. Abstract: A fundamental task in multi-agent systems is to match agents to alternatives (e.g., resources or tasks). Often, this is accomplished by eliciting agents' ordinal rankings over the alternatives instead of their exact numerical utilities.
Reservoir Network with Structural Plasticity for Human Activity Recognition
Zyarah, Abdullah M., Abdul-Hadi, Alaa M., Kudithipudi, Dhireesha
--The unprecedented dissemination of edge devices is accompanied by a growing demand for neuromorphic chips that can process time-series data natively without cloud support. Echo state network (ESN) is a class of recurrent neural networks that can be used to identify unique patterns in time-series data and predict future events. It is known for minimal computing resource requirements and fast training, owing to the use of linear optimization solely at the readout stage. In this work, a custom-design neuromorphic chip based on ESN targeting edge devices is proposed. The proposed system supports various learning mechanisms, including structural plasticity and synaptic plasticity, locally on-chip. This provides the network with an additional degree of freedom to continuously learn, adapt, and alter its structure and sparsity level, ensuring high performance and continuous stability. We demonstrate the performance of the proposed system as well as its robustness to noise against real-world time-series datasets while considering various topologies of data movement. An average accuracy of 95.95% and 85.24% are achieved on human activity recognition and prosthetic finger control, respectively. HE last decade has seen significant advancement in neuromorphic computing with a major thrust centered around processing streaming data using recurrent neural networks (RNNs). Despite the fact RNNs demonstrate promising performance in numerous domains including speech recognition [1], computer vision [2], stock trading [3], and medical diagnosis [4], such networks suffer from slow convergence and intensive computations [5]. In order to bypass these challenges, Jaeger and Maass suggest leveraging the rich dynamics offered by the networks' recurrent connections and random parameters and limit the training to the network advanced layers, particularly the readout layer [7]-[9]. With that, the network training and its computation complexity are significantly simplified. There are three classes of RNN networks trained using this approach known as a liquid state machine (LSM) [7], delayed-feedback reservoir [10], [11], and echo state network (ESN) which is going to be the focus of this work. ESN is demonstrated in a variety of tasks, including pattern recognition, anomaly detection [12], spatial-temporal forecasting [13], and modeling dynamic motions in bio-mimic robots [14].
Engadget Podcast: iPhone 16e review and Amazon's AI-powered Alexa
The keyword for the iPhone 16e seems to be "compromise." In this episode, Devindra chats with Cherlynn about her iPhone 16e review and try to figure out who this phone is actually for. Also, they dive into Amazon's Alexa event, where we finally learned more about the company's AI-powered voice assistant. Alexa seems useful, but can we trust it? Listen below or subscribe on your podcast app of choice. If you've got suggestions or topics you'd like covered on the show, be sure to email us or drop a note in the comments! And be sure to check out our other podcast, Engadget News! Framework unveils a cheap 2-in-1 laptop and aโฆmodular desktop? Devindra: This week, it's the iPhone 16e, which Cherlynn has reviewed. We're going to get her full thoughts on that thing. And also, Amazon held an AI event this week. We expected a lot of devices, but they spent 75 minutes talking about Alexa plus, which is the AI powered Alexa. Cherlynn: we expected a lot of devices. Cherlynn: one, at least one it's been a while. Devindra: Mr. Panos Panay was there, the father of the service and no devices, just him talking about AI. Cherlynn: Oh, and stay tuned at the end of this episode. Uh, I, we included an interview that I did with, um, the vice president of Alexa to talk more about the new Alexa plus. Devindra: Anyway, folks, if you're enjoying the show, please be sure to subscribe to us on iTunes or your podcaster of choice, leave us a review on iTunes and drop us an email at podcast@engadget.com. You can also join us on our live [00:01:00] stream on Thursday mornings, typically around 11 a. m. Um, you'll see our faces. Sometimes we'll do Q& A and show off devices as well. This week, uh, Sherilyn has the iPhone 16e, which is the least, um, impressive thing to show off. It's just like, Hey, you have an iPhone from 10 years ago, five, a while ago, Devindra: last, was there a single camera back iPhone? Cherlynn: Oh God, before that was 11. So, you know, it's like a flashback. So let's talk about this thing, Sherlynn. And I checked out your review. First of all, you gave it a really, um, I think serviceable score. Your title is what's your acceptable compromise. And really when we were talking about it last week, it really was like compromise seemed like the key word. The thing we kept coming back to was like just one camera, no mag safe, no fast wireless charging. What are your overall thoughts on this thing? Cherlynn: I mean, so that headline is like all thanks to our EIC, Aaron [00:02:00]Souppouris, because I was like, where, where do I go from here? How do I, so, so he's right. It is like, instead of what's in your wallet, it's like, what are you willing to take out your wallet? I'll tell you the story. So yesterday I was at the Amazon devices and services event where there were no devices and A bunch of other reporters had gathered and we were all like, you know, the, like, review's going up soon, right?
A Ukrainian Family's Three Years of War
One morning last month, while I was waiting at a bus stop on the western edge of the western Ukrainian city of Lviv, I struck up a conversation with a man in his early forties named Mykola Hryhoryan. Across from the bus stop was a bombed-out museum. I asked if he knew what had happened to it. "It was hit by a Russian drone," he said. Mykola was wearing jeans and a black parka with the hood pulled over his head. He told me that he was a soldier.
Narwhals spotted using tusks for non-mating fun
With their long, spiral tusks, narwhals (Monodon monoceros) look like something out of a fairy tale. Primarily seen in male narwhals, these single elongated teeth that can grow up to 10 feet. These gregarious whales typically travel in pods of two to 10 individuals, but are a bit elusive and difficult to study in the wild. Scientists believe that the tusks are primarily used in competition for mates, but that might not be the whole story. New drone evidence detailed in a study published February 28 in the journal Frontiers in Marine Science found that narwhals can use their tusks to forage, explore their surroundings, and even play.
Congratulations to the #AAAI2025 award winners
A number of prestigious AAAI awards were presented during the official opening ceremony of the Thirty-Ninth AAAI Conference on Artificial Intelligence (AAAI 2025) on 27 February. Some of the winners will also be giving invited talks as part of the programme. The AAAI Award for Artificial Intelligence for Humanity recognises the positive impacts of artificial intelligence to protect, enhance, and improve human life in meaningful ways with long-lived effects. The winner of this year's award is Stuart J. Russell (University of California, Berkeley, USA). Stuart has been recognised for "work on the conceptual and theoretical foundations of provably beneficial AI and his leadership in creating the field of AI safety".
Interview with AAAI Fellow Sriraam Natarajan: Human-allied AI
Each year the AAAI recognizes a group of individuals who have made significant, sustained contributions to the field of artificial intelligence by appointing them as Fellows. Over the course of the next few months, we'll be talking to some of the 2025 AAAI Fellows. In this interview we hear from Sriraam Natarajan, Professor at the University of Texas at Dallas, who was elected as a Fellow for "significant contributions to statistical relational AI, healthcare adaptations and service to the AAAI community". We find out about his career path, research on human-allied AI, reflections on changes to the AI landscape, and passion for cricket. Could you start by telling us about your career so far, where you work and your broad area of research?
MEBench: Benchmarking Large Language Models for Cross-Document Multi-Entity Question Answering
Multi-entity question answering (MEQA) represents significant challenges for large language models (LLM) and retrieval-augmented generation (RAG) systems, which frequently struggle to consolidate scattered information across diverse documents. While existing methods excel at single-document comprehension, they often struggle with cross-document aggregation, particularly when resolving entity-dense questions like "What is the distribution of ACM Fellows among various fields of study?", which require integrating entity-centric insights from heterogeneous sources (e.g., Wikipedia pages). To address this gap, we introduce MEBench, a novel multi-document, multi-entity benchmark designed to systematically evaluate LLMs' capacity to retrieve, consolidate, and reason over fragmented information. Our benchmark comprises 4,780 questions which are systematically categorized into three primary categories, further divided into eight distinct types, ensuring broad coverage of real-world multi-entity reasoning scenarios. Our experiments on state-of-the-art LLMs (e.g., GPT-4, Llama-3) and RAG pipelines reveal critical limitations: even advanced models achieve only 59% accuracy on MEBench. Our benchmark emphasizes the importance of completeness and factual precision of information extraction in MEQA tasks, using Entity-Attributed F1 (EA-F1) metric for granular evaluation of entity-level correctness and attribution validity. MEBench not only highlights systemic weaknesses in current LLM frameworks but also provides a foundation for advancing robust, entity-aware QA architectures.