Goto

Collaborating Authors

 Agents


Bridging Visualization and Optimization: Multimodal Large Language Models on Graph-Structured Combinatorial Optimization

arXiv.org Artificial Intelligence

Graph-structured combinatorial challenges are inherently difficult due to their nonlinear and intricate nature, often rendering traditional computational methods ineffective or expensive. However, these challenges can be more naturally tackled by humans through visual representations that harness our innate ability for spatial reasoning. In this study, we propose transforming graphs into images to preserve their higher-order structural features accurately, revolutionizing the representation used in solving graph-structured combinatorial tasks. This approach allows machines to emulate human-like processing in addressing complex combinatorial challenges. By combining the innovative paradigm powered by multimodal large language models (MLLMs) with simple search techniques, we aim to develop a novel and effective framework for tackling such problems. Our investigation into MLLMs spanned a variety of graph-based tasks, from combinatorial problems like influence maximization to sequential decision-making in network dismantling, as well as addressing six fundamental graph-related issues. Our findings demonstrate that MLLMs exhibit exceptional spatial intelligence and a distinctive capability for handling these problems, significantly advancing the potential for machines to comprehend and analyze graph-structured data with a depth and intuition akin to human cognition. These results also imply that integrating MLLMs with simple optimization strategies could form a novel and efficient approach for navigating graph-structured combinatorial challenges without complex derivations, computationally demanding training and fine-tuning.


Navigating Robot Swarm Through a Virtual Tube with Flow-Adaptive Distribution Control

arXiv.org Artificial Intelligence

With the rapid development of robot swarm technology and its diverse applications, navigating robot swarms through complex environments has emerged as a critical research direction. To ensure safe navigation and avoid potential collisions with obstacles, the concept of virtual tubes has been introduced to define safe and navigable regions. However, current control methods in virtual tubes face the congestion issues, particularly in narrow virtual tubes with low throughput. To address these challenges, we first originally introduce the concepts of virtual tube area and flow capacity, and develop an new evolution model for the spatial density function. Next, we propose a novel control method that combines a modified artificial potential field (APF) for swarm navigation and density feedback control for distribution regulation, under which a saturated velocity command is designed. Then, we generate a global velocity field that not only ensures collision-free navigation through the virtual tube, but also achieves locally input-to-state stability (LISS) for density tracking errors, both of which are rigorously proven. Finally, numerical simulations and realistic applications validate the effectiveness and advantages of the proposed method in managing robot swarms within narrow virtual tubes.


Reviews: Learning to Communicate with Deep Multi-Agent Reinforcement Learning

Neural Information Processing Systems

Though the paper contains a very thorough experimental evaluation of the suggested DIAL technique for multi-agent settings and the reviewer understands that it would have taken a lot of time and effort to set up and evaluate the experiments, the paper does not make a very novel contribution. It is clear that idea of shared memory and passing message gradients between agents would speed up learning and help to find the optimal policy faster, but this might not be a very natural way to do it. For instance, humans working in teams do not have shared memories. Also, for humans, messages from other humans are a part of their observation at each time step, rather than separate signals which are treated specially as messages and optimized differently than the rest of the observation. The idea of passing message gradients is certainly useful to have trainable message protocols while training a set of agents to perform a repetitive task, but doesn't offer much insight or useful interpretation as to how humans perform tasks in teams.


Reviews: Long-term Causal Effects via Behavioral Game Theory

Neural Information Processing Systems

Typically (in the Rubin potential outcomes model, which is what you are building on), the causal effect is defined at the individual level, with a "treatment" outcome and "control" outcome for each experimental unit. The fundamental problem of causal inference is that only one of these two outcomes is actually observed for each experimental unit. You seem to be focusing on a slightly different issue, which is that the effect of treating the entire population cannot be determined correctly from just data when half the population is treated. It seems to me that this issue -- which can arise due to a variety of violations of the SUTVA assumption -- can exist independent of whether there is a multiagent interaction. Conversely, it seems multiagent considerations are relevant even when defining causal effects at the sub-population level.


Reviews: Learning Multiagent Communication with Backpropagation

Neural Information Processing Systems

The model is a deep network which consists of a stack of layers, with parameter sharing between modules of a same layer. This parameter sharing allows the number of agents to vary during the task. Also, it allows to drastically reduce the number of parameters to be learned. The key idea of the paper is to use the output of every module of a given layer to build the communication input for the next layer. While this appears to obtain interesting results in the reported experiments, I find this proposal very straightforward and poorly innovative, as it corresponds to a quiet classical neural network structure.


Reviews: Bi-Objective Online Matching and Submodular Allocations

Neural Information Processing Systems

The paper builds on the following problem. There is a set of agents. Each agent a is associated a submodular function f_a. Elements from the universe arrive over time. When an element arrives, it is to be assigned to an agent.


Multi-Agent Learning with Heterogeneous Linear Contextual Bandits

Neural Information Processing Systems

As trained intelligent systems become increasingly pervasive, multiagent learning has emerged as a popular framework for studying complex interactions between autonomous agents. Yet, a formal understanding of how and when learners in heterogeneous environments benefit from sharing their respective experiences is far from complete. In this paper, we seek answers to these questions in the context of linear contextual bandits. We present a novel distributed learning algorithm based on the upper confidence bound (UCB) algorithm, which we refer to as H-LINUCB, wherein agents cooperatively minimize the group regret under the coordination of a central server. In the setting where the level of heterogeneity or dissimilarity across the environments is known to the agents, we show that H-LINUCB is provably optimal in regimes where the tasks are highly similar or highly dissimilar.


Decompose a Task into Generalizable Subtasks in Multi-Agent Reinforcement Learning

Neural Information Processing Systems

In recent years, Multi-Agent Reinforcement Learning (MARL) techniques have made significant strides in achieving high asymptotic performance in single task. However, there has been limited exploration of model transferability across tasks. Training a model from scratch for each task can be time-consuming and expensive, especially for large-scale Multi-Agent Systems. Therefore, it is crucial to develop methods for generalizing the model across tasks. Considering that there exist task-independent subtasks across MARL tasks, a model that can decompose such subtasks from the source task could generalize to target tasks.


The Best of Both Worlds in Network Population Games: Reaching Consensus and Convergence to Equilibrium

Neural Information Processing Systems

Reaching consensus and convergence to equilibrium are two major challenges of multi-agent systems. Although each has attracted significant attention, relatively few studies address both challenges at the same time. This paper examines the connection between the notions of consensus and equilibrium in a multi-agent system where multiple interacting sub-populations coexist. We argue that consensus can be seen as an intricate component of intra-population stability, whereas equilibrium can be seen as encoding inter-population stability. We show that smooth fictitious play, a well-known learning model in game theory, can achieve both consensus and convergence to equilibrium in diverse multi-agent settings.


DIFFER:Decomposing Individual Reward for Fair Experience Replay in Multi-Agent Reinforcement Learning

Neural Information Processing Systems

Cooperative multi-agent reinforcement learning (MARL) is a challenging task, as agents must learn complex and diverse individual strategies from a shared team reward. However, existing methods struggle to distinguish and exploit important individual experiences, as they lack an effective way to decompose the team reward into individual rewards. To address this challenge, we propose DIFFER, a powerful theoretical framework for decomposing individual rewards to enable fair experience replay in MARL.By enforcing the invariance of network gradients, we establish a partial differential equation whose solution yields the underlying individual reward function. The individual TD-error can then be computed from the solved closed-form individual rewards, indicating the importance of each piece of experience in the learning task and guiding the training process. Our method elegantly achieves an equivalence to the original learning framework when individual experiences are homogeneous, while also adapting to achieve more muscular efficiency and fairness when diversity is observed.Our extensive experiments on popular benchmarks validate the effectiveness of our theory and method, demonstrating significant improvements in learning efficiency and fairness. Code is available in supplement material.