Agents
ABC: Adversarial Behavioral Cloning for Offline Mode-Seeking Imitation Learning
Hudson, Eddy, Durugkar, Ishan, Warnell, Garrett, Stone, Peter
Given a dataset of expert agent interactions with an environment of interest, a viable method to extract an effective agent policy is to estimate the maximum likelihood policy indicated by this data. This approach is commonly referred to as behavioral cloning (BC). In this work, we describe a key disadvantage of BC that arises due to the maximum likelihood objective function; namely that BC is mean-seeking with respect to the state-conditional expert action distribution when the learner's policy is represented with a Gaussian. To address this issue, we introduce a modified version of BC, Adversarial Behavioral Cloning (ABC), that exhibits mode-seeking behavior by incorporating elements of GAN (generative adversarial network) training. We evaluate ABC on toy domains and a domain based on Hopper from the DeepMind Control suite, and show that it outperforms standard BC by being mode-seeking in nature.
Coordination with Humans via Strategy Matching
Zhao, Michelle, Simmons, Reid, Admoni, Henny
Human and robot partners increasingly need to work together to perform tasks as a team. Robots designed for such collaboration must reason about how their task-completion strategies interplay with the behavior and skills of their human team members as they coordinate on achieving joint goals. Our goal in this work is to develop a computational framework for robot adaptation to human partners in human-robot team collaborations. We first present an algorithm for autonomously recognizing available task-completion strategies by observing human-human teams performing a collaborative task. By transforming team actions into low dimensional representations using hidden Markov models, we can identify strategies without prior knowledge. Robot policies are learned on each of the identified strategies to construct a Mixture-of-Experts model that adapts to the task strategies of unseen human partners. We evaluate our model on a collaborative cooking task using an Overcooked simulator. Results of an online user study with 125 participants demonstrate that our framework improves the task performance and collaborative fluency of human-agent teams, as compared to state of the art reinforcement learning methods.
Developing Decentralised Resilience to Malicious Influence in Collective Perception Problem
Wise, Chris, Hussein, Aya, El-Fiqi, Heba
In collective decision-making, designing algorithms that use only local information to effect swarm-level behaviour is a non-trivial problem. We used machine learning techniques to teach swarm members to map their local perceptions of the environment to an optimal action. A curriculum inspired by Machine Education approaches was designed to facilitate this learning process and teach the members the skills required for optimal performance in the collective perception problem. We extended upon previous approaches by creating a curriculum that taught agents resilience to malicious influence. The experimental results show that well-designed rules-based algorithms can produce effective agents. When performing opinion fusion, we implemented decentralised resilience by having agents dynamically weight received opinion. We found a non-significant difference between constant and dynamic weights, suggesting that momentum-based opinion fusion is perhaps already a resilience mechanism.
Motion Style Transfer: Modular Low-Rank Adaptation for Deep Motion Forecasting
Kothari, Parth, Li, Danya, Liu, Yuejiang, Alahi, Alexandre
Motion forecasting is an essential pillar for the successful deployment of autonomous systems in environments comprising various heterogeneous agents. It presents the challenges of modeling (i) universal etiquette (e.g., goal-directed behaviors, avoiding collisions) that govern general motion dynamics of all agents; and (ii) social norms (e.g., the minimum separation distance, preferred speed) that influence the navigation styles of different agents across different locations. Owing to the success of deep neural networks on large-scale datasets, learning prediction models in a data-driven manner has become a de-facto approach for motion forecasting and has shown impressive results [1, 2, 3, 4]. However, existing deep forecasting models suffer from inferior performance when they encounter novel scenarios [5, 6, 7, 8]. For instance, a network trained with large-scale data for pedestrian forecasting struggles to directly generalize to cyclists. Some recent methods propose to incorporate strong priors robust to the underlying distribution shifts [9, 10, 11]. Yet, these priors often make strong assumptions on the distribution shifts, which may not hold in practice.
Decentralized Policy Optimization
The study of decentralized learning or independent learning in cooperative multi-agent reinforcement learning has a history of decades. Recently empirical studies show that independent PPO (IPPO) can obtain good performance, close to or even better than the methods of centralized training with decentralized execution, in several benchmarks. However, decentralized actor-critic with convergence guarantee is still open. In this paper, we propose \textit{decentralized policy optimization} (DPO), a decentralized actor-critic algorithm with monotonic improvement and convergence guarantee. We derive a novel decentralized surrogate for policy optimization such that the monotonic improvement of joint policy can be guaranteed by each agent \textit{independently} optimizing the surrogate. In practice, this decentralized surrogate can be realized by two adaptive coefficients for policy optimization at each agent. Empirically, we compare DPO with IPPO in a variety of cooperative multi-agent tasks, covering discrete and continuous action spaces, and fully and partially observable environments. The results show DPO outperforms IPPO in most tasks, which can be the evidence for our theoretical results.
Collaborative Video Analytics on Distributed Edges with Multiagent Deep Reinforcement Learning
Dong, Yuqi, Gao, Guanyu, Wang, Ran, Yan, Zhisheng
Deep Neural Network (DNN) based video analytics empowers many computer vision-based applications to achieve high recognition accuracy. To reduce inference delay and bandwidth cost for video analytics, the DNN models can be deployed on the edge nodes, which are proximal to end users. However, the processing capacity of an edge node is limited, potentially incurring substantial delay if the inference requests on an edge node is overloaded. While efforts have been made to enhance video analytics by optimizing the configurations on a single edge node, we observe that multiple edge nodes can work collaboratively by utilizing the idle resources on each other to improve the overall processing capacity and resource utilization. To this end, we propose a Multiagent Reinforcement Learning (MARL) based approach, named as EdgeVision, for collaborative video analytics on distributed edges. The edge nodes can jointly learn the optimal policies for video preprocessing, model selection, and request dispatching by collaborating with each other to minimize the overall cost. We design an actor-critic-based MARL algorithm with an attention mechanism to learn the optimal policies. We build a multi-edge-node testbed and conduct experiments with real-world datasets to evaluate the performance of our method. The experimental results show our method can improve the overall rewards by 33.6%-86.4% compared with the most competitive baseline methods.
SRIBO: An Efficient and Resilient Single-Range and Inertia Based Odometry for Flying Robots
Dong, Wei, Mei, Zheyuan, Ying, Yuanjiong, Chen, Sijia, ie, Yichen, Zhu, Xiangyang
Positioning with one inertial measurement unit and one ranging sensor is commonly thought to be feasible only when trajectories are in certain patterns ensuring observability. For this reason, to pursue observable patterns, it is required either exciting the trajectory or searching key nodes in a long interval, which is commonly highly nonlinear and may also lack resilience. Therefore, such a positioning approach is still not widely accepted in real-world applications. To address this issue, this work first investigates the dissipative nature of flying robots considering aerial drag effects and re-formulates the corresponding positioning problem, which guarantees observability almost surely. On this basis, a dimension-reduced wriggling estimator is proposed accordingly. This estimator slides the estimation horizon in a stepping manner, and output matrices can be approximately evaluated based on the historical estimation sequence. The computational complexity is then further reduced via a dimension-reduction approach using polynomial fittings. In this way, the states of robots can be estimated via linear programming in a sufficiently long interval, and the degree of observability is thereby further enhanced because an adequate redundancy of measurements is available for each estimation. Subsequently, the estimator's convergence and numerical stability are proven theoretically. Finally, both indoor and outdoor experiments verify that the proposed estimator can achieve decimeter-level precision at hundreds of hertz per second, and it is resilient to sensors' failures. Hopefully, this study can provide a new practical approach for self-localization as well as relative positioning of cooperative agents with low-cost and lightweight sensors.
Graph Reinforcement Learning Application to Co-operative Decision-Making in Mixed Autonomy Traffic: Framework, Survey, and Challenges
Liu, Qi, Li, Xueyuan, Li, Zirui, Wu, Jingda, Du, Guodong, Gao, Xin, Yang, Fan, Yuan, Shihua
Proper functioning of connected and automated vehicles (CAVs) is crucial for the safety and efficiency of future intelligent transport systems. Meanwhile, transitioning to fully autonomous driving requires a long period of mixed autonomy traffic, including both CAVs and human-driven vehicles. Thus, collaboration decision-making for CAVs is essential to generate appropriate driving behaviors to enhance the safety and efficiency of mixed autonomy traffic. In recent years, deep reinforcement learning (DRL) has been widely used in solving decision-making problems. However, the existing DRL-based methods have been mainly focused on solving the decision-making of a single CAV. Using the existing DRL-based methods in mixed autonomy traffic cannot accurately represent the mutual effects of vehicles and model dynamic traffic environments. To address these shortcomings, this article proposes a graph reinforcement learning (GRL) approach for multi-agent decision-making of CAVs in mixed autonomy traffic. First, a generic and modular GRL framework is designed. Then, a systematic review of DRL and GRL methods is presented, focusing on the problems addressed in recent research. Moreover, a comparative study on different GRL methods is further proposed based on the designed framework to verify the effectiveness of GRL methods. Results show that the GRL methods can well optimize the performance of multi-agent decision-making for CAVs in mixed autonomy traffic compared to the DRL methods. Finally, challenges and future research directions are summarized. This study can provide a valuable research reference for solving the multi-agent decision-making problems of CAVs in mixed autonomy traffic and can promote the implementation of GRL-based methods into intelligent transportation systems. The source code of our work can be found at https://github.com/Jacklinkk/Graph_CAVs.
HeRoSwarm: Fully-Capable Miniature Swarm Robot Hardware Design With Open-Source ROS Support
Starks, Michael, Gupta, Aryan, Venkata, Sanjay Sarma Oruganti, Parasuraman, Ramviyas
Experiments using large numbers of miniature swarm robots are desirable to teach, study, and test multi-robot and swarm intelligence algorithms and their applications. To realize the full potential of a swarm robot, it should be capable of not only motion but also sensing, computing, communication, and power management modules with multiple options. Current swarm robot platforms developed for commercial and academic research purposes lack several of these critical attributes by focusing only on a few of these aspects. Therefore, in this paper, we propose the HeRoSwarm, a fully-capable swarm robot platform with open-source hardware and software support. The proposed robot hardware is a low-cost design with commercial off-the-shelf components that uniquely integrates multiple sensing, communication, and computing modalities with various power management capabilities into a tiny footprint. Moreover, our swarm robot with odometry capability with Robot Operating Systems (ROS) support is unique in its kind. This simple yet powerful swarm robot design has been extensively verified with different prototyping variants and multi-robot experimental demonstrations.
A Multi-Criteria Metaheuristic Algorithm for Distributed Optimization of Electric Energy Storage
Schrage, Rico, Tiemann, Paul Hendrik, Nieße, Astrid
The distributed schedule optimization of energy storage constitutes a challenge. Such algorithms often expect an input set containing all feasible schedules or respectively require to efficiently search the schedule space. It is hardly possible to accomplish this with energy storage due to its high flexibility. In this paper, the problem is introduced in detail and addressed by a metaheuristic algorithm, which generates a preselection of schedules. Three contributions are presented to achieve this goal: First, an extension for a distributed schedule optimization allowing a simultaneous optimization is developed. Second, an evolutionary algorithm is designed to generate optimized schedules. Third, the algorithm is extended to include an arbitrary local criterion. It is shown that the presented approach is suitable to schedule electric energy storage in real households and industries with different generator and storage types.