Goto

Collaborating Authors

 Agents


Learning to Learn Group Alignment: A Self-Tuning Credo Framework with Multiagent Teams

arXiv.org Artificial Intelligence

Mixed incentives among a population with multiagent teams has been shown to have advantages over a fully cooperative system; however, discovering the best mixture of incentives or team structure is a difficult and dynamic problem. We propose a framework where individual learning agents self-regulate their configuration of incentives through various parts of their reward function. This work extends previous work by giving agents the ability to dynamically update their group alignment during learning and by allowing teammates to have different group alignment. Our model builds on ideas from hierarchical reinforcement learning and meta-learning to learn the configuration of a reward function that supports the development of a behavioral policy. We provide preliminary results in a commonly studied multiagent environment and find that agents can achieve better global outcomes by self-tuning their respective group alignment parameters.


Exact Subspace Diffusion for Decentralized Multitask Learning

arXiv.org Artificial Intelligence

Classical paradigms for distributed learning, such as federated or decentralized gradient descent, employ consensus mechanisms to enforce homogeneity among agents. While these strategies have proven effective in i.i.d. scenarios, they can result in significant performance degradation when agents follow heterogeneous objectives or data. Distributed strategies for multitask learning, on the other hand, induce relationships between agents in a more nuanced manner, and encourage collaboration without enforcing consensus. We develop a generalization of the exact diffusion algorithm for subspace constrained multitask learning over networks, and derive an accurate expression for its mean-squared deviation when utilizing noisy gradient approximations. We verify numerically the accuracy of the predicted performance expressions, as well as the improved performance of the proposed approach over alternatives based on approximate projections.


A Dynamic Heterogeneous Team-based Non-iterative Approach for Online Pick-up and Just-In-Time Delivery Problems

arXiv.org Artificial Intelligence

This paper presents a non-iterative approach for finding the assignment of heterogeneous robots to efficiently execute online Pickup and Just-In-Time Delivery (PJITD) tasks with optimal resource utilization. The PJITD assignments problem is formulated as a spatio-temporal multi-task assignment (STMTA) problem. The physical constraints on the map and vehicle dynamics are incorporated in the cost formulation. The linear sum assignment problem is formulated for the heterogeneous STMTA problem. The recently proposed Dynamic Resource Allocation with Multi-task assignments (DREAM) approach has been modified to solve the heterogeneous PJITD problem. At the start, it computes the minimum number of robots required (with their types) to execute given heterogeneous PJITD tasks. These required robots are added to the team to guarantee the feasibility of all PJITD tasks. Then robots in an updated team are assigned to execute the PJITD tasks while minimizing the total cost for the team to execute all PJITD tasks. The performance of the proposed non-iterative approach has been validated using high-fidelity software-in-loop simulations and hardware experiments. The simulations and experimental results clearly indicate that the proposed approach is scalable and provides optimal resource utilization.


Optimal and Efficient Auctions for the Gradual Procurement of Strategic Service Provider Agents

Journal of Artificial Intelligence Research

We consider an outsourcing problem where a software agent procures multiple services  from providers with uncertain reliabilities to complete a computational task before a  strict deadline. The service consumer’s goal is to design an outsourcing strategy (defining  which services to procure and when) so as to maximize a specific objective function. This  objective function can be different based on the consumer’s nature; a socially-focused consumer  often aims to maximize social welfare, while a self-interested consumer often aims  to maximize its own utility. However, in both cases, the objective function depends on  the providers’ execution costs, which are privately held by the self-interested providers and  hence may be misreported to influence the consumer’s decisions. For such settings, we  develop a unified approach to design truthful procurement auctions that can be used by  both socially-focused and, separately, self-interested consumers. This approach benefits  from our proposed weighted threshold payment scheme which pays the provably minimum  amount to make an auction with a monotone outsourcing strategy incentive compatible.  This payment scheme can handle contingent outsourcing plans, where additional procurement  happens gradually over time and only if the success probability of the already hired  providers drops below a time-dependent threshold. Using a weighted threshold payment  scheme, we design two procurement auctions that maximize, as well as two low-complexity  heuristic-based auctions that approximately maximize, the consumer’s expected utility and  expected social welfare, respectively. We demonstrate the effectiveness and strength of our  proposed auctions through both game-theoretical and empirical analysis. 


Generative Agents: Stanford's Groundbreaking AI Study Simulates Authentic Human Behavior - Artisana

#artificialintelligence

A new study by a team of Stanford AI researchers introduces a groundbreaking concept: Generative Agents, computer programs that employ generative models to simulate authentic human behavior. The innovative architecture developed by the researchers enables these agents to demonstrate human-like abilities in memory storage and retrieval, introspection on motivations and goals, and planning and reacting to novel situations. Powered by ChatGPT, these agents engage with each other and researchers as if they were genuine human beings. In the experiment, the researchers placed 25 generative agents within a virtual world resembling a sandbox video game, similar to The Sims. Each agent was assigned a unique background and participated in a two-day simulation.


Forget Chatbots, The Future is Autonomous Agents

#artificialintelligence

Sci-fi has been predicting autonomous systems for decades. One example is Tony Stark's J.A.R.V.I.S. from the Avengers: Tony Stark: "J.A.R.V.I.S., are you up?" Tony: "I'd like to open a new project file, index as: Mark II." J: "Shall I store this on the Stark Industries' central database?" Tony: "I don't know who to trust right now. 'Til further notice, why don't we just keep everything on my private server." J: "Working on a secret project, are we, sir?" Tony: "I don't want this winding up in the wrong hands. Maybe in mine, it could actually do some good."


A model of communication-enabled traffic interactions

arXiv.org Artificial Intelligence

A major challenge for autonomous vehicles is handling interactive scenarios, such as highway merging, with human-driven vehicles. A better understanding of human interactive behaviour could help address this challenge. Such understanding could be obtained through modelling human behaviour. However, existing modelling approaches predominantly neglect communication between drivers and assume that some drivers in the interaction only respond to others, but do not actively influence them. Here we argue that addressing these two limitations is crucial for accurate modelling of interactions. We propose a new computational framework addressing these limitations. Similar to game-theoretic approaches, we model the interaction in an integral way rather than modelling an isolated driver who only responds to their environment. Contrary to game theory, our framework explicitly incorporates communication and bounded rationality. We demonstrate the model in a simplified merging scenario, illustrating that it generates plausible interactive behaviour (e.g., aggressive and conservative merging). Furthermore, human-like gap-keeping behaviour emerged in a car-following scenario directly from risk perception without the explicit implementation of time or distance gaps in the model's decision-making. These results suggest that our framework is a promising approach to interaction modelling that can support the development of interaction-aware autonomous vehicles.


Multi-Layer Continuum Deformation Optimization of Multi-Agent Systems

arXiv.org Artificial Intelligence

This paper studies the problem of safe and optimal continuum deformation of a large-scale multi-agent system (MAS). We present a novel approach for MAS continuum deformation coordination that aims to achieve safe and efficient agent movement using a leader-follower multi-layer hierarchical optimization framework with a single input layer, multiple hidden layers, and a single output layer. The input layer receives the reference (material) positions of the primary leaders, the hidden layers compute the desired positions of the interior leader agents and followers, and the output layer computes the nominal position of the MAS configuration. By introducing a lower bound on the major principles of the strain field of the MAS deformation, we obtain linear inequality safety constraints and ensure inter-agent collision avoidance. The continuum deformation optimization is formulated as a quadratic programming problem. It consists of the following components: (i) decision variables that represent the weights in the first hidden layer; (ii) a quadratic cost function that penalizes deviation of the nominal MAS trajectory from the desired MAS trajectory; and (iii) inequality safety constraints that ensure inter-agent collision avoidance. To validate the proposed approach, we simulate and present the results of continuum deformation on a large-scale quadcopter team tracking a desired helix trajectory, demonstrating improvements in safety and efficiency.


Model-based Dynamic Shielding for Safe and Efficient Multi-Agent Reinforcement Learning

arXiv.org Artificial Intelligence

Multi-Agent Reinforcement Learning (MARL) discovers policies that maximize reward but do not have safety guarantees during the learning and deployment phases. Although shielding with Linear Temporal Logic (LTL) is a promising formal method to ensure safety in single-agent Reinforcement Learning (RL), it results in conservative behaviors when scaling to multi-agent scenarios. Additionally, it poses computational challenges for synthesizing shields in complex multi-agent environments. This work introduces Model-based Dynamic Shielding (MBDS) to support MARL algorithm design. Our algorithm synthesizes distributive shields, which are reactive systems running in parallel with each MARL agent, to monitor and rectify unsafe behaviors. The shields can dynamically split, merge, and recompute based on agents' states. This design enables efficient synthesis of shields to monitor agents in complex environments without coordination overheads. We also propose an algorithm to synthesize shields without prior knowledge of the dynamics model. The proposed algorithm obtains an approximate world model by interacting with the environment during the early stage of exploration, making our MBDS enjoy formal safety guarantees with high probability. We demonstrate in simulations that our framework can surpass existing baselines in terms of safety guarantees and learning performance.


SA-reCBS: Multi-robot task assignment with integrated reactive path generation

arXiv.org Artificial Intelligence

Yifan Bai, Christoforos Kanellakis and George Nikolakopoulos Robotics and AI Team Luleå University of Technology, Sweden Abstract: In this paper, we study the multi-robot task assignment and path-finding problem (MRTAPF), where a number of robots are required to visit all given tasks while avoiding collisions with each other. We propose a novel two-layer algorithm SA-reCBS that cascades the simulated annealing algorithm and conflict-based search to solve this problem. Compared to other approaches in the field of MRTAPF, the advantage of SA-reCBS is that without requiring a pre-bundle of tasks to groups with the same number of groups as the number of robots, it enables a part of robots needed to visit all tasks in collision-free paths. We test the algorithm in various simulation instances and compare it with state-of-the-art algorithms. The result shows that SA-reCBS has a better performance with a higher success rate, less computational time, and better objective values.