Agents
Artificial virtuous agents in a multiagent tragedy of the commons
Although virtue ethics has repeatedly been proposed as a suitable framework for the development of artificial moral agents (AMAs), it has been proven difficult to approach from a computational perspective. In this work, we present the first technical implementation of artificial virtuous agents (AVAs) in moral simulations. First, we review previous conceptual and technical work in artificial virtue ethics and describe a functionalistic path to AVAs based on dispositional virtues, bottom-up learning, and top-down eudaimonic reward. We then provide the details of a technical implementation in a moral simulation based on a tragedy of the commons scenario. The experimental results show how the AVAs learn to tackle cooperation problems while exhibiting core features of their theoretical counterpart, including moral character, dispositional virtues, learning from experience, and the pursuit of eudaimonia. Ultimately, we argue that virtue ethics provides a compelling path toward morally excellent machines and that our work provides an important starting point for such endeavors.
Moving Virtual Agents Forward in Space and Time
Silva, Gabriel F., Knob, Paulo, Johansson, Carlos G., Schlatter, Douglas A., Musse, Soraia R.
This article proposes an adaptation from the model of Bianco for fast-forwarding agents in crowd simulation, which enables us to accurately fast forward agents in time. Besides being able to jump from one position to another, agents are able to stay inside their track, it means, the new position is calculated taking into account the original global path the agent would follow, if not being fast-forwarded. Obstacles and other agents around are also taken into account when calculating the new position. In addition, we included a personality aspect on agents, which affect their behaviors and, also, be taken into account when jumping to a future time and space. We conducted some experiments to validate our model, which shows that it was able to indeed fast forward agents from a position to another, in a coherent time, sticking to a given global path while avoiding collisions. Finally, we present a use case, showing that our method can fit inside a "Fog of War" system.
Learning Algorithms for Intelligent Agents and Mechanisms
In this thesis, we research learning algorithms for optimal decision making in two different contexts, Reinforcement Learning in Part I and Auction Design in Part II. Reinforcement learning (RL) is an area of machine learning that is concerned with how an agent should act in an environment in order to maximize its cumulative reward over time. In Chapter 2, inspired by statistical physics, we develop a novel approach to Reinforcement Learning (RL) that not only learns optimal policies with enhanced desirable properties but also sheds new light on maximum entropy RL. In Chapter 3, we tackle the generalization problem in RL using a Bayesian perspective. We show that imperfect knowledge of the environments dynamics effectively turn a fully-observed Markov Decision Process (MDP) into a Partially Observed MDP (POMDP) that we call the Epistemic POMDP. Informed by this observation, we develop a new policy learning algorithm LEEP which has improved generalization properties. Designing an incentive compatible, individually rational auction that maximizes revenue is a challenging and intractable problem. Recently, deep learning based approaches have been proposed to learn optimal auctions from data. While successful, this approach suffers from a few limitations, including sample inefficiency, lack of generalization to new auctions, and training difficulties. In Chapter 4, we construct a symmetry preserving neural network architecture, EquivariantNet, suitable for anonymous auctions. EquivariantNet is not only more sample efficient but is also able to learn auction rules that generalize well to other settings. In Chapter 5, we propose a novel formulation of the auction learning problem as a two player game. The resulting learning algorithm, ALGNet, is easier to train, more reliable and better suited for non stationary settings.
Astrobotics: Swarm Robotics for Astrophysical Studies
Macktoobian, Matin, Gillet, Denis, Kneib, Jean-Paul
Published in "IEEE Robotics and Automation Magazine", DOI: 10.1109/MRA.2020.3044911 Matin Macktoobian, Denis Gillet, and Jean-Paul Kneib The authors are with the School of Engineering, Swiss Federal Institute of Technology in Lausanne (EPFL), Lausanne, Switzerland (e-mail: matin.macktoobian@epfl.ch; Abstract This paper introduces the emerging field of astrobotics, that is, a recently-established branch of robotics to be of service to astrophysics and observational astronomy. We first describe a modern requirement of dark matter studies, i.e., the generation of the map of the observable universe, using astrobots. Astrobots differ from conventional two-degree-of-freedom robotic manipulators in two respects. First, the dense formation of astrobots give rise to the extremely overlapping dynamics of neighboring astrobots which make them severely subject to collisions. Second, the structure of astrobots and their mechanical specifications are specialized due to the embedded optical fibers passed through them. We focus on the coordination problem of astrobots whose solutions shall be collision-free, fast execution, and complete in terms of the astrobots' convergence rates. We also illustrate the significant impact of astrobots assignments to observational targets on the quality of coordination solutions To present the current state of the field, we elaborate the open problems including next-generation astrophysical projects including 20,000 astrobots, and other fields, such as space debris tracking, in which astrobots may be potentially used. Astrobotics is an emerging field of swarm robotics aiming to the development and control of astrobots [1, 2] to be of service to astrophysical studies and cosmological spectroscopic observations. In particular, astrobotics addresses a wide range of swarm-robotic-related topics (see, Figure 1) which exhibit challenging problems in design, interaction, coordination, and mission planning corresponding to astrobots. There have been many astrophysical projects, such as the SDSS family [3] which seek the generation of the map of the observable universe.
Exploring Effectiveness of Explanations for Appropriate Trust: Lessons from Cognitive Psychology
Verhagen, Ruben S., Mehrotra, Siddharth, Neerincx, Mark A., Jonker, Catholijn M., Tielman, Myrthe L.
The rapid development of Artificial Intelligence (AI) requires developers and designers of AI systems to focus on the collaboration between humans and machines. AI explanations of system behavior and reasoning are vital for effective collaboration by fostering appropriate trust, ensuring understanding, and addressing issues of fairness and bias. However, various contextual and subjective factors can influence an AI system explanation's effectiveness. This work draws inspiration from findings in cognitive psychology to understand how effective explanations can be designed. We identify four components to which explanation designers can pay special attention: perception, semantics, intent, and user & context. We illustrate the use of these four explanation components with an example of estimating food calories by combining text with visuals, probabilities with exemplars, and intent communication with both user and context in mind. We propose that the significant challenge for effective AI explanations is an additional step between explanation generation using algorithms not producing interpretable explanations and explanation communication. We believe this extra step will benefit from carefully considering the four explanation components outlined in our work, which can positively affect the explanation's effectiveness.
Game Theoretic Rating in N-player general-sum games with Equilibria
Marris, Luke, Lanctot, Marc, Gemp, Ian, Omidshafiei, Shayegan, McAleer, Stephen, Connor, Jerome, Tuyls, Karl, Graepel, Thore
Rating strategies in a game is an important area of research in game theory and artificial intelligence, and can be applied to any real-world competitive or cooperative setting. Traditionally, only transitive dependencies between strategies have been used to rate strategies (e.g. Elo), however recent work has expanded ratings to utilize game theoretic solutions to better rate strategies in non-transitive games. This work generalizes these ideas and proposes novel algorithms suitable for N-player, general-sum rating of strategies in normal-form games according to the payoff rating system. This enables well-established solution concepts, such as equilibria, to be leveraged to efficiently rate strategies in games with complex strategic interactions, which arise in multiagent training and real-world interactions between many agents. We empirically validate our methods on real world normal-form data (Premier League) and multiagent reinforcement learning agent evaluation.
Cost Aware Asynchronous Multi-Agent Active Search
Banerjee, Arundhati, Ghods, Ramina, Schneider, Jeff
Multi-agent active search requires autonomous agents to choose sensing actions that efficiently locate targets. In a realistic setting, agents also must consider the costs that their decisions incur. Previously proposed active search algorithms simplify the problem by ignoring uncertainty in the agent's environment, using myopic decision making, and/or overlooking costs. In this paper, we introduce an online active search algorithm to detect targets in an unknown environment by making adaptive cost-aware decisions regarding the agent's actions. Our algorithm combines principles from Thompson Sampling (for search space exploration and decentralized multi-agent decision making), Monte Carlo Tree Search (for long horizon planning) and pareto-optimal confidence bounds (for multi-objective optimization in an unknown environment) to propose an online lookahead planner that removes all the simplifications. We analyze the algorithm's performance in simulation to show its efficacy in cost aware active search.
From Intelligent Agents to Trustworthy Human-Centred Multiagent Systems
Soorati, Mohammad Divband, Gerding, Enrico H., Marchioni, Enrico, Naumov, Pavel, Norman, Timothy J., Ramchurn, Sarvapali D., Rastegari, Bahar, Sobey, Adam, Stein, Sebastian, Tarpore, Danesh, Yazdanpanah, Vahid, Zhang, Jie
The Agents, Interaction and Complexity research group at the University of Southampton has a long track record of research in multiagent systems (MAS). We have made substantial scientific contributions across learning in MAS, game-theoretic techniques for coordinating agent systems, and formal methods for representation and reasoning. We highlight key results achieved by the group and elaborate on recent work and open research challenges in developing trustworthy autonomous systems and deploying human-centred AI systems that aim to support societal good.
INTERACT: Achieving Low Sample and Communication Complexities in Decentralized Bilevel Learning over Networks
Liu, Zhuqing, Zhang, Xin, Khanduri, Prashant, Lu, Songtao, Liu, Jia
In recent years, decentralized bilevel optimization problems have received In recent years, fueled by the rise of machine learning and artificial increasing attention in the networking and machine learning intelligence in edge networks, decentralized bilevel optimization communities thanks to their versatility in modeling decentralized problems have received increasing attention in the networking and learning problems over peer-to-peer networks (e.g., multi-agent machine learning communities. This is due to the versatility of meta-learning, multi-agent reinforcement learning, personalized decentralized bilevel optimization in supporting many decentralized training, and Byzantine-resilient learning). However, for decentralized learning paradigms over peer-to-peer networks, such as the bilevel optimization over peer-to-peer networks with limited multi-agent versions of meta learning [22, 33, 33], hyperparameter computation and communication capabilities, how to achieve low optimization problem[24, 29], area under curve (AUC) problems sample and communication complexities are two fundamental challenges [19, 32], and reinforcement learning[9, 40]. To date, however, that remain under-explored so far. In this paper, we make there remain many challenges and open problems in decentralized the first attempt to investigate the class of decentralized bilevel bilevel learning over peer-to-peer networks. Two of the most optimization problems with nonconvex and strongly-convex structure fundamental challenges in decentralized bilevel optimization are corresponding to the outer and inner subproblems, respectively.
PlaneSDF-based Change Detection for Long-term Dense Mapping
Fu, Jiahui, Lin, Chengyuan, Taguchi, Yuichi, Cohen, Andrea, Zhang, Yifu, Mylabathula, Stephen, Leonard, John J.
The ability to process environment maps across multiple sessions is critical for robots operating over extended periods of time. Specifically, it is desirable for autonomous agents to detect changes amongst maps of different sessions so as to gain a conflict-free understanding of the current environment. In this paper, we look into the problem of change detection based on a novel map representation, dubbed Plane Signed Distance Fields (PlaneSDF), where dense maps are represented as a collection of planes and their associated geometric components in SDF volumes. Given point clouds of the source and target scenes, we propose a three-step PlaneSDF-based change detection approach: (1) PlaneSDF volumes are instantiated within each scene and registered across scenes using plane poses; 2D height maps and object maps are extracted per volume via height projection and connected component analysis. (2) Height maps are compared and intersected with the object map to produce a 2D change location mask for changed object candidates in the source scene. (3) 3D geometric validation is performed using SDF-derived features per object candidate for change mask refinement. We evaluate our approach on both synthetic and real-world datasets and demonstrate its effectiveness via the task of changed object detection. Supplementary video: https://youtu.be/oh-MQPWTwZI