Agents
MMFN: Multi-Modal-Fusion-Net for End-to-End Driving
Zhang, Qingwen, Tang, Mingkai, Geng, Ruoyu, Chen, Feiyi, Xin, Ren, Wang, Lujia
Inspired by the fact that humans use diverse sensory organs to perceive the world, sensors with different modalities are deployed in end-to-end driving to obtain the global context of the 3D scene. In previous works, camera and LiDAR inputs are fused through transformers for better driving performance. These inputs are normally further interpreted as high-level map information to assist navigation tasks. Nevertheless, extracting useful information from the complex map input is challenging, for redundant information may mislead the agent and negatively affect driving performance. We propose a novel approach to efficiently extract features from vectorized High-Definition (HD) maps and utilize them in the end-to-end driving tasks. In addition, we design a new expert to further enhance the model performance by considering multi-road rules. Experimental results prove that both of the proposed improvements enable our agent to achieve superior performance compared with other methods.
Zero-Shot Style Transfer for Gesture Animation driven by Text and Speech using Adversarial Disentanglement of Multimodal Style Encoding
Fares, Mireille, Grimaldi, Michele, Pelachaud, Catherine, Obin, Nicolas
Modeling virtual agents with behavior style is one factor for personalizing human-agent interaction. In this paper, we propose an efficient yet effective machine learning approach to synthesize gestures driven by prosodic features and text in the style of different speakers including those unseen during training. Our model performs zero-shot multimodal style transfer driven by multimodal data from the PATS database containing videos of various speakers. We view style as being pervasive while speaking; it colors the communicative behaviors expressivity while speech content is carried by multimodal signals and text. This disentanglement scheme of content and style allows us to directly infer the style embedding even of speaker whose data are not part of the training phase, without requiring any further training or fine-tuning. The first goal of our model is to generate the gestures of a source speaker based on the content of two input modalities - Mel spectrogram and text semantics. The second goal is to condition the source speaker's predicted gestures on the multimodal behavior style embedding of a target speaker. The third goal is to allow zero-shot style transfer of speakers unseen during training without re-training the model. Our system consists of two main components: (1) a speaker style encoder network that learns to generate a fixed-dimensional speaker embedding style from a target speaker multimodal data (mel-spectrogram, pose, and text); and (2) a sequence-to-sequence synthesis network that synthesizes gestures based on the content of the input modalities - text and mel-spectrogram - of a source speaker, and conditioned on the speaker style embedding. We evaluate that our model is able to synthesize gestures of a source speaker given the two input modalities, and transfer the knowledge of target speaker style variability learned by the speaker style encoder to the gesture generation task in a zero-shot setup, indicating that the model has learned a high quality speaker representation. For our evaluation we convert the 2D generated gestures to 3D poses, and produce 3D animations of the generated gestures. We conduct objective and subjective evaluations to validate our approach and compare it with baselines. Keywords: audio and text driven gesture synthesis, zero-shot style transfer, embodied conversational agents 1 INTRODUCTION Human behavior style is a socially meaningful clustering of features found within and across multiple modalities, specifically in linguistic [7], spoken behavior such as the speaking style conveyed by speech prosody [29, 33], and nonverbal behavior such as hand gestures and body posture [32, 42].
A Game-Theoretic Approach for Hierarchical Epidemic Control
Jia, Feiran, Mate, Aditya, Li, Zun, Jabbari, Shahin, Chakraborty, Mithun, Tambe, Milind, Wellman, Michael, Vorobeychik, Yevgeniy
Democratic governments and institutions typically have a hierarchical structure. For example, policies in the U.S., Canada, and many European democracies emerge from complex interactions among the federal and state governments, as well as county boards, city councils and mayors. Such interactions are characterized by inherent asymmetries across different levels of the hierarchy. On the one hand, the specifics of policy formulation and enforcement (e.g., training and deployment of personnel and updating of infrastructure) are generally in the hands of administrative bodies at lower levels of the hierarchy -- often the lowest level -- for practical reasons; actions these entities take are the ones that truly matter in the sense that they directly impact costs and benefits realized at all levels. On the other hand, entities at higher levels may have the power to impose constraints in some form or another on the policy-makers within their immediate jurisdiction (e.g., the U.S. federal government can constrain state policies); violations of these constraints, in turn, entail a noncompliance cost to the violator, such as legal costs, penalties, or reputation loss. Examples of such hierarchical policy structure arise in the spheres of education (e.g., topics to be included in primary education), healthcare (e.g., vaccination) and immigration. A preeminent recent example of such hierarchical policy-making is the response to the ongoing COVID-19 pandemic in countries with decentralized administration. Policies concerning social distancing, masking and vaccination have involved recommendations at the federal level, guidelines and restrictions at the state/province/district level, and measures adopted by specific counties, cities or even individual businesses and schools. In general, policies are contentious.
Evaluating Inter-Operator Cooperation Scenarios to Save Radio Access Network Energy
Marjou, Xavier, Gléau, Tangui Le, Messié, Vincent, Radier, Benoit, Lemlouma, Tayeb, Fromentoux, Gaël
Reducing energy consumption is crucial to reduce the human debt's with regard to our planet. Therefore most companies try to reduce their energetic consumption while taking care to preserve the service delivered to their customers. To do so, a service provider (SP) typically downscale or shutdown part of its infrastructure in periods of low-activity where only few customers need the service. However an SP still needs to maintain part of its infrastructure "on", which still requires significant energy. For example a mobile national operator (MNO) needs to maintain most of its radio access network (RAN) active. Could an SP do better by cooperating with other SPs who would temporarily support its users, thus allowing it to temporarily shut down its infrastructure, and then reciprocate during another low-activity period? To answer this question, we investigated a novel collaboration framework based on multi-agent reinforcement learning (MARL) allowing negotiations between SPs as well as trustful reports from a distributed ledger technology (DLT) to evaluate the amount of energy being saved. We leveraged it to experiment three different sets of rules (free, recommended, or imposed) regulating the negotiation between multiple SPs (3, 4, 8, or 10). With respect to four cooperation metrics (efficiency, safety, incentive-compatibility, and fairness), the simulations showed that the imposed set of rules proved to be the best mode.
Asymptotic Tracking Control of Uncertain MIMO Nonlinear Systems with Less Conservative Controllability Conditions
Zhou, Bing, Huang, Xiucai, Song, Yongduan
For uncertain multiple inputs multi-outputs (MIMO) nonlinear systems, it is nontrivial to achieve asymptotic tracking, and most existing methods normally demand certain controllability conditions that are rather restrictive or even impractical if unexpected actuator faults are involved. In this note, we present a method capable of achieving zero-error steady-state tracking with less conservative (more practical) controllability condition. By incorporating a novel Nussbaum gain technique and some positive integrable function into the control design, we develop a robust adaptive asymptotic tracking control scheme for the system with time-varying control gain being unknown its magnitude and direction. By resorting to the existence of some feasible auxiliary matrix, the current state-of-art controllability condition is further relaxed, which enlarges the class of systems that can be considered in the proposed control scheme. All the closed-loop signals are ensured to be globally ultimately uniformly bounded. Moreover, such control methodology is further extended to the case involving intermittent actuator faults, with application to robotic systems. Finally, simulation studies are carried out to demonstrate the effectiveness and flexibility of this method.
Deep Reinforcement Learning for Multi-Agent Interaction
Ahmed, Ibrahim H., Brewitt, Cillian, Carlucho, Ignacio, Christianos, Filippos, Dunion, Mhairi, Fosong, Elliot, Garcin, Samuel, Guo, Shangmin, Gyevnar, Balint, McInroe, Trevor, Papoudakis, Georgios, Rahman, Arrasy, Schäfer, Lukas, Tamborski, Massimiliano, Vecchio, Giuseppe, Wang, Cheng, Albrecht, Stefano V.
The development of autonomous agents which can interact with other agents to accomplish a given task is a core area of research in artificial intelligence and machine learning. Towards this goal, the Autonomous Agents Research Group develops novel machine learning algorithms for autonomous systems control, with a specific focus on deep reinforcement learning and multi-agent reinforcement learning. Research problems include scalable learning of coordinated agent policies and inter-agent communication; reasoning about the behaviours, goals, and composition of other agents from limited observations; and sample-efficient learning based on intrinsic motivation, curriculum learning, causal inference, and representation learning. This article provides a broad overview of the ongoing research portfolio of the group and discusses open problems for future directions.
Decentralized Learning With Limited Communications for Multi-robot Coverage of Unknown Spatial Fields
Nakamura, Kensuke, Santos, María, Leonard, Naomi Ehrich
This paper presents an algorithm for a team of mobile robots to simultaneously learn a spatial field over a domain and spatially distribute themselves to optimally cover it. Drawing from previous approaches that estimate the spatial field through a centralized Gaussian process, this work leverages the spatial structure of the coverage problem and presents a decentralized strategy where samples are aggregated locally by establishing communications through the boundaries of a Voronoi partition. We present an algorithm whereby each robot runs a local Gaussian process calculated from its own measurements and those provided by its Voronoi neighbors, which are incorporated into the individual robot's Gaussian process only if they provide sufficiently novel information. The performance of the algorithm is evaluated in simulation and compared with centralized approaches.
Heterogeneous-Agent Mirror Learning: A Continuum of Solutions to Cooperative MARL
Kuba, Jakub Grudzien, Feng, Xidong, Ding, Shiyao, Dong, Hao, Wang, Jun, Yang, Yaodong
The necessity for cooperation among intelligent machines has popularised cooperative multi-agent reinforcement learning (MARL) in the artificial intelligence (AI) research community. However, many research endeavours have been focused on developing practical MARL algorithms whose effectiveness has been studied only empirically, thereby lacking theoretical guarantees. As recent studies have revealed, MARL methods often achieve performance that is unstable in terms of reward monotonicity or suboptimal at convergence. To resolve these issues, in this paper, we introduce a novel framework named Heterogeneous-Agent Mirror Learning (HAML) that provides a general template for MARL algorithmic designs. We prove that algorithms derived from the HAML template satisfy the desired properties of the monotonic improvement of the joint reward and the convergence to Nash equilibrium. We verify the practicality of HAML by proving that the current state-of-the-art cooperative MARL algorithms, HATRPO and HAPPO, are in fact HAML instances. Next, as a natural outcome of our theory, we propose HAML extensions of two well-known RL algorithms, HAA2C (for A2C) and HADDPG (for DDPG), and demonstrate their effectiveness against strong baselines on StarCraftII and Multi-Agent MuJoCo tasks.
An Introduction to Multi-Agent Reinforcement Learning and Review of its Application to Autonomous Mobility
Schmidt, Lukas M., Brosig, Johanna, Plinge, Axel, Eskofier, Bjoern M., Mutschler, Christopher
Many scenarios in mobility and traffic involve multiple different agents that need to cooperate to find a joint solution. Recent advances in behavioral planning use Reinforcement Learning to find effective and performant behavior strategies. However, as autonomous vehicles and vehicle-to-X communications become more mature, solutions that only utilize single, independent agents leave potential performance gains on the road. Multi-Agent Reinforcement Learning (MARL) is a research field that aims to find optimal solutions for multiple agents that interact with each other. This work aims to give an overview of the field to researchers in autonomous mobility. We first explain MARL and introduce important concepts. Then, we discuss the central paradigms that underlie MARL algorithms, and give an overview of state-of-the-art methods and ideas in each paradigm. With this background, we survey applications of MARL in autonomous mobility scenarios and give an overview of existing scenarios and implementations.
Decomposing your complex AI problem: Hierarchy
Problem worlds often come with an innate hierarchy. Naturally, this may prompt the question: which level(s) of the hierarchy should be modelled? For example, the US Stock Market can be modelled as a whole or at the index level -- think, the Dow Jones, or for individual stocks. In a linear system, the way that the lower levels interact with the upper levels is "linear" or directly correlated. Take the example of an analytics system for business intelligence and reporting -- sales, inventories, etc.