Goto

Collaborating Authors

 Agents


Sea-cret Agents: Maritime Abduction for Region Generation to Expose Dark Vessel Trajectories

arXiv.org Artificial Intelligence

Bad actors in the maritime industry engage in illegal behaviors after disabling their vessel's automatic identification system (AIS) - which makes finding such vessels difficult for analysts. Machine learning approaches only succeed in identifying the locations of these ``dark vessels'' in the immediate future. This work leverages ideas from the literature on abductive inference applied to locating adversarial agents to solve the problem. Specifically, we combine concepts from abduction, logic programming, and rule learning to create an efficient method that approaches full recall of dark vessels while requiring less search area than machine learning methods. We provide a logic-based paradigm for reasoning about maritime vessels, an abductive inference query method, an automatically extracted rule-based behavior model methodology, and a thorough suite of experiments.


Autotelic Reinforcement Learning: Exploring Intrinsic Motivations for Skill Acquisition in Open-Ended Environments

arXiv.org Artificial Intelligence

Intelligence, which leverages sociocultural interactions to enhance open-ended skill acquisition. Artificial Intelligence (AI) aims to create autonomous agents that can operate across diverse environments and complete a wide range of tasks. Researchers pursue different approaches, each focusing on specific drivers of learning. In Reinforcement Learning (RL) [1], agents learn by exploring their environment and using their experience to solve tasks. Imitation Learning (IL) [2] involves agents learning from expert demonstrations, while Multi-Agent Reinforcement Learning (MARL) [3] emphasizes cooperation among agents to solve collaborative tasks. Recent advancements in RL have demonstrated success in varied domains, such as playing Atari games [4], mastering chess and Go [5], and controlling stratospheric balloons [6]. IL, combined with transformers [7], has enabled generalist agents to be trained on diverse datasets and to perform in-context reinforcement learning via algorithm distillation. However, these algorithms remain sample-inefficient and struggle with generalization, creativity, and tackling novel tasks, largely because they rely on isolated learning signals. This research explores sociocultural interactions as a new avenue for AI learning inspired by human development.


Simulating the Emergence of Differential Case Marking with Communicating Neural-Network Agents

arXiv.org Artificial Intelligence

Differential Case Marking (DCM) refers to the phenomenon where grammatical case marking is applied selectively based on semantic, pragmatic, or other factors. The emergence of DCM has been studied in artificial language learning experiments with human participants, which were specifically aimed at disentangling the effects of learning from those of communication (Smith & Culbertson, 2020). Multi-agent reinforcement learning frameworks based on neural networks have gained significant interest to simulate the emergence of human-like linguistic phenomena. In this study, we employ such a framework in which agents first acquire an artificial language before engaging in communicative interactions, enabling direct comparisons to human result. Using a very generic communication optimization algorithm and neural-network learners that have no prior experience with language or semantic preferences, our results demonstrate that learning alone does not lead to DCM, but when agents communicate, differential use of markers arises. This supports Smith and Culbertson (2020)'s findings that highlight the critical role of communication in shaping DCM and showcases the potential of neural-agent models to complement experimental research on language evolution.


Online Learning of Counter Categories and Ratings in PvP Games

arXiv.org Artificial Intelligence

In competitive games, strength ratings like Elo are widely used to quantify player skill and support matchmaking by accounting for skill disparities better than simple win rate statistics. However, scalar ratings cannot handle complex intransitive relationships, such as counter strategies seen in Rock-Paper-Scissors. To address this, recent work introduced Neural Rating Table and Neural Counter Table, which combine scalar ratings with discrete counter categories to model intransitivity. While effective, these methods rely on neural network training and cannot perform real-time updates. In this paper, we propose an online update algorithm that extends Elo principles to incorporate real-time learning of counter categories. Our method dynamically adjusts both ratings and counter relationships after each match, preserving the explainability of scalar ratings while addressing intransitivity. Experiments on zero-sum competitive games demonstrate its practicality, particularly in scenarios without complex team compositions.


Fully Autonomous AI Agents Should Not be Developed

arXiv.org Artificial Intelligence

This paper argues that fully autonomous AI agents should not be developed. In support of this position, we build from prior scientific literature and current product marketing to delineate different AI agent levels and detail the ethical values at play in each, documenting trade-offs in potential benefits and risks. Our analysis reveals that risks to people increase with the autonomy of a system: The more control a user cedes to an AI agent, the more risks to people arise. Particularly concerning are safety risks, which affect human life and impact further values.


DECAF: Learning to be Fair in Multi-agent Resource Allocation

arXiv.org Artificial Intelligence

A wide variety of resource allocation problems operate under resource constraints that are managed by a central arbitrator, with agents who evaluate and communicate preferences over these resources. We formulate this broad class of problems as Distributed Evaluation, Centralized Allocation (DECA) problems and propose methods to learn fair and efficient policies in centralized resource allocation. Our methods are applied to learning long-term fairness in a novel and general framework for fairness in multi-agent systems. We show three different methods based on Double Deep Q-Learning: (1) A joint weighted optimization of fairness and utility, (2) a split optimization, learning two separate Q-estimators for utility and fairness, and (3) an online policy perturbation to guide existing black-box utility functions toward fair solutions. Our methods outperform existing fair MARL approaches on multiple resource allocation domains, even when evaluated using diverse fairness functions, and allow for flexible online trade-offs between utility and fairness.


Enhancing Online Learning Efficiency Through Heterogeneous Resource Integration with a Multi-Agent RAG System

arXiv.org Artificial Intelligence

However, navigating and synthesizing information across these disparate sources can be a timeintensive Efficient online learning requires seamless access to diverse resources and inefficient process, creating barriers to efficient online such as videos, code repositories, documentation, and general learning [8]. The challenges associated with multi-source learning web content. This poster paper introduces early-stage work are especially evident in technical domains, where the need to on a Multi-Agent Retrieval-Augmented Generation (RAG) System quickly find accurate and relevant information is critical. For instance, designed to enhance learning efficiency by integrating these heterogeneous a developer exploring a new framework might consult a resources. Using specialized agents tailored for specific YouTube tutorial for an overview, reference a GitHub repository resource types (e.g., YouTube tutorials, GitHub repositories, documentation for implementation details, examine the official documentation for websites, and search engines), the system automates deeper insights, and conduct general web searches for troubleshooting.


Constant-Factor Distortion Mechanisms for $k$-Committee Election

arXiv.org Artificial Intelligence

In the $k$-committee election problem, we wish to aggregate the preferences of $n$ agents over a set of alternatives and select a committee of $k$ alternatives that minimizes the cost incurred by the agents. While we typically assume that agent preferences are captured by a cardinal utility function, in many contexts we only have access to ordinal information, namely the agents' rankings over the outcomes. As preference rankings are not as expressive as cardinal utilities, a loss of efficiency is inevitable, and is quantified by the notion of \emph{distortion}. We study the problem of electing a $k$-committee that minimizes the sum of the $\ell$-largest costs incurred by the agents, when agents and candidates are embedded in a metric space. This problem is called the $\ell$-centrum problem and captures both the utilitarian and egalitarian objectives. When $k \geq 2$, it is not possible to compute a bounded-distortion committee using purely ordinal information. We develop the first algorithms (that we call mechanisms) for the $\ell$-centrum problem (when $k \geq 2$), which achieve $O(1)$-distortion while eliciting only a very limited amount of cardinal information via value queries. We obtain two types of query-complexity guarantees: $O(\log k \log n)$ queries \emph{per agent}, and $O(k^2 \log^2 n)$ queries \emph{in total} (while achieving $O(1)$-distortion in both cases). En route, we give a simple adaptive-sampling algorithm for the $\ell$-centrum $k$-clustering problem.


Multi-agent Architecture Search via Agentic Supernet

arXiv.org Artificial Intelligence

Large Language Model (LLM)-empowered multi-agent systems extend the cognitive boundaries of individual agents through disciplined collaboration and interaction, while constructing these systems often requires labor-intensive manual designs. Despite the availability of methods to automate the design of agentic workflows, they typically seek to identify a static, complex, one-size-fits-all system, which, however, fails to dynamically allocate inference resources based on the difficulty and domain of each query. To address this challenge, we shift away from the pursuit of a monolithic agentic system, instead optimizing the \textbf{agentic supernet}, a probabilistic and continuous distribution of agentic architectures. We introduce MaAS, an automated framework that samples query-dependent agentic systems from the supernet, delivering high-quality solutions and tailored resource allocation (\textit{e.g.}, LLM calls, tool calls, token cost). Comprehensive evaluation across six benchmarks demonstrates that MaAS \textbf{(I)} requires only $6\sim45\%$ of the inference costs of existing handcrafted or automated multi-agent systems, \textbf{(II)} surpasses them by $0.54\%\sim11.82\%$, and \textbf{(III)} enjoys superior cross-dataset and cross-LLM-backbone transferability.


Deep Meta Coordination Graphs for Multi-agent Reinforcement Learning

arXiv.org Artificial Intelligence

This paper presents deep meta coordination graphs (DMCG) for learning cooperative policies in multi-agent reinforcement learning (MARL). Coordination graph formulations encode local interactions and accordingly factorize the joint value function of all agents to improve efficiency in MARL. However, existing approaches rely solely on pairwise relations between agents, which potentially oversimplifies complex multi-agent interactions. DMCG goes beyond these simple direct interactions by also capturing useful higher-order and indirect relationships among agents. It generates novel graph structures accommodating multiple types of interactions and arbitrary lengths of multi-hop connections in coordination graphs to model such interactions. It then employs a graph convolutional network module to learn powerful representations in an end-to-end manner. We demonstrate its effectiveness in multiple coordination problems in MARL where other state-of-the-art methods can suffer from sample inefficiency or fail entirely. All codes can be found here: https://github.com/Nikunj-Gupta/dmcg-marl.