Agents
Inverse Online Learning: Understanding Non-Stationary and Reactionary Policies
Chan, Alex J., Curth, Alicia, van der Schaar, Mihaela
Human decision making is well known to be imperfect and the ability to analyse such processes individually is crucial when attempting to aid or improve a decision-maker's ability to perform a task, e.g. to alert them to potential biases or oversights on their part. To do so, it is necessary to develop interpretable representations of how agents make decisions and how this process changes over time as the agent learns online in reaction to the accrued experience. To then understand the decision-making processes underlying a set of observed trajectories, we cast the policy inference problem as the inverse to this online learning problem. By interpreting actions within a potential outcomes framework, we introduce a meaningful mapping based on agents choosing an action they believe to have the greatest treatment effect. We introduce a practical algorithm for retrospectively estimating such perceived effects, alongside the process through which agents update them, using a novel architecture built upon an expressive family of deep state-space models. Through application to the analysis of UNOS organ donation acceptance decisions, we demonstrate that our approach can bring valuable insights into the factors that govern decision processes and how they change over time.
MUG: Interactive Multimodal Grounding on User Interfaces
Li, Tao, Li, Gang, Zheng, Jingjie, Wang, Purple, Li, Yang
We present MUG, a novel interactive task for multimodal grounding where a user and an agent work collaboratively on an interface screen. Prior works modeled multimodal UI grounding in one round: the user gives a command and the agent responds to the command. Yet, in a realistic scenario, a user command can be ambiguous when the target action is inherently difficult to articulate in natural language. MUG allows multiple rounds of interactions such that upon seeing the agent responses, the user can give further commands for the agent to refine or even correct its actions. Such interaction is critical for improving grounding performances in real-world use cases. To investigate the problem, we create a new dataset that consists of 77,820 sequences of human user-agent interaction on mobile interfaces in which 20% involves multiple rounds of interactions. To establish our benchmark, we experiment with a range of modeling variants and evaluation strategies, including both offline and online evaluation-the online strategy consists of both human evaluation and automatic with simulators. Our experiments show that allowing iterative interaction significantly improves the absolute task completion by 18% over the entire test dataset and 31% over the challenging subset. Our results lay the foundation for further investigation of the problem.
Multi-Agent Path Finding: A New Boolean Encoding
Asín Achá, Roberto (Universidad de Concepción) | López, Rodrigo (Universidad de Chile & Pontificia Universidad Católica de Chile) | Hagedorn, Sebastian (Pontificia Universidad Católica de Chile) | Baier, Jorge A. (Pontificia Universidad Católica de Chile)
Multi-agent pathfinding (MAPF) is an NP-hard problem. As such, dense maps may be very hard to solve optimally. In such scenarios, compilation-based approaches, via Boolean satisfiability (SAT) and answer set programming (ASP), have been shown to outperform heuristic-search-based approaches, such as conflict-based search (CBS). In this paper, we propose a new Boolean encoding for MAPF, and show how to implement it in ASP and MaxSAT. A feature that distinguishes our encoding from existing ones is that swap and follow conflicts are encoded using binary clauses, which can be exploited by current conflict-driven clause learning (CDCL) solvers. In addition, the number of clauses used to encode swap and follow conflicts do not depend on the number of agents, allowing us to scale better. For MaxSAT, we study different ways in which we may combine the MSU3 and LSU algorithms for maximum performance. In our experimental evaluation, we used square grids, ranging from 20 x 20 to 50 x 50 cells, and warehouse maps, with a varying number of agents and obstacles. We compared against representative solvers of the state-of-the-art, including the search-based algorithm CBS, the ASP-based solver ASP-MAPF, and the branch-and-cut-and-price hybrid solver, BCP. We observe that the ASP implementation of our encoding, ASP-MAPF2 outperforms other solvers in most of our experiments. The MaxSAT implementation of our encoding, MtMS shows best performance in relatively small warehouse maps when the number of agents is large, which are the instances with closer resemblance to hard puzzle-like problems.
Hierarchical Integration of Model Predictive and Fuzzy Logic Control for Combined Coverage and Target-Oriented Search-and-Rescue via Robots with Imperfect Sensors
de Koning, Christopher, Jamshidnejad, Anahita
Search-and-rescue (SaR) in unknown environments requires precise, optimal, and fast decisions. Robots are promising candidates for autonomously performing SaR tasks in unknown environments. While humans use their heuristics to effectively deal with uncertainties, optimisation of multiple objectives in the presence of physical and control constraints is a mathematical challenge that requires machine computations. Thus having both human-inspired and mathematical control capabilities is desired for SaR robots. Moreover, coordinating the decisions of robots with little computation cost in large-scale SaR missions is an open challenge. Finally, in real-life data perceived by SaR robots may be prone to uncertainties. We introduce a hierarchical multi-agent control architecture that exploits non-homogeneous and imperfect perception capabilities of SaR robots, as well as the computational efficiency and robustness to failure of decentralised control methods and global performance improvement of centralised control methods. The integrated structure of the proposed control framework allows to combine human-inspired and mathematical decision making methods in a coordinated and computationally efficient way. The results of various computer-based simulations show that while the area coverage of the proposed approach is comparable to existing heuristic methods that are particularly developed for coverage-oriented SaR, the efficiency of the introduced approach in locating the trapped victims is significantly higher. Furthermore, with comparable computation times, the proposed control approach successfully avoids potential conflicts that exist in non-cooperative methods. These results confirm that the proposed multi-agent control system is capable of combining coverage-oriented and target-oriented SaR in a balanced and coordinated way.
Reinforcement Learning with Tensor Networks: Application to Dynamical Large Deviations
Gillman, Edward, Rose, Dominic C., Garrahan, Juan P.
We present a framework to integrate tensor network (TN) methods with reinforcement learning (RL) for solving dynamical optimisation tasks. We consider the RL actor-critic method, a model-free approach for solving RL problems, and introduce TNs as the approximators for its policy and value functions. Our "actor-critic with tensor networks" (ACTeN) method is especially well suited to problems with large and factorisable state and action spaces. As an illustration of the applicability of ACTeN we solve the exponentially hard task of sampling rare trajectories in two paradigmatic stochastic models, the East model of glasses and the asymmetric simple exclusion process (ASEP), the latter being particularly challenging to other methods due to the absence of detailed balance. With substantial potential for further integration with the vast array of existing RL methods, the approach introduced here is promising both for applications in physics and to multi-agent RL problems more generally.
SkiNet, A Petri Net Generation Tool for the Verification of Skillset-based Autonomous Systems
Pelletier, Baptiste, Lesire, Charles, Doose, David, Godary-Dejean, Karen, Dramé-Maigné, Charles
The need for high-level autonomy and robustness of autonomous systems for missions in dynamic and remote environment has pushed developers to come up with new software architectures. A common architecture style is to summarize the capabilities of the robotic system into elementary actions, called skills, on top of which a skill management layer is implemented to structure, test and control the functional layer. However, current available verification tools only provide either mission-specific verification or verification on a model that does not replicate the actual execution of the system, which makes it difficult to ensure its robustness to unexpected events. To that end, a tool, SkiNet, has been developed to transform the skill-based architecture of a system into a Petri net modeling the state-machine behaviors of the skills and the resources they handle. The Petri net allows the use of model-checking, such as Linear Temporal Logic (LTL) or Computational Tree Logic (CTL), for the user to analyze and verify the model of the system.
DMCA: Dense Multi-agent Navigation using Attention and Communication
Arul, Senthil Hariharan, Bedi, Amrit Singh, Manocha, Dinesh
In decentralized multi-robot navigation, the agents lack the world knowledge to make safe and (near-)optimal plans reliably and make their decisions on their neighbors' observable states. We present a reinforcement learning based multi-agent navigation algorithm that performs inter-agent communications. In order to deal with the variable number of neighbors for each agent, we use a multi-head self-attention mechanism to encode neighbor information and create a fixed-length observation vector. We pose communication selection as a link prediction problem, where the network predicts whether communication is necessary given the observable information. The communicated information augments the observed neighbor information and is used to select a suitable navigation plan. We highlight the benefits of our approach by performing safe and efficient navigation among multiple robots in dense and challenging benchmarks. We also compare the performance with other learning-based methods and highlight improvements in terms of fewer collisions and time-to-goal in dense scenarios.
Advising Autonomous Cars about the Rules of the Road
Collenette, Joe, Dennis, Louise A., Fisher, Michael
This paper describes (R)ules (o)f (T)he (R)oad (A)dvisor, an agent that provides recommended and possible actions to be generated from a set of human-level rules. We describe the architecture and design of RoTRA, both formally and with an example. Specifically, we use RoTRA to formalise and implement the UK "Rules of the Road", and describe how this can be incorporated into autonomous cars such that they can reason internally about obeying the rules of the road. In addition, the possible actions generated are annotated to indicate whether the rules state that the action must be taken or that they only recommend that the action should be taken, as per the UK Highway Code (Rules of The Road). The benefits of utilising this system include being able to adapt to different regulations in different jurisdictions; allowing clear traceability from rules to behaviour, and providing an external automated accountability mechanism that can check whether the rules were obeyed in some given situation. A simulation of an autonomous car shows, via a concrete example, how trust can be built by putting the autonomous vehicle through a number of scenarios which test the car's ability to obey the rules of the road. Autonomous cars that incorporate this system are able to ensure that they are obeying the rules of the road and external (legal or regulatory) bodies can verify that this is the case, without the vehicle or its manufacturer having to expose their source code or make their working transparent, thus allowing greater trust between car companies, jurisdictions, and the general public.
Verifying Safety of Behaviour Trees in Event-B
Tadiello, Matteo, Troubitsyna, Elena
Autonomous Systems (AS) like Humanoid Robots, Autonomous Vehicles, or Unmanned Aerial Vehicles are becoming increasingly complex and need to interact with dynamic environments and with each other. For this reason, robots require tools to enable advanced perception and understanding of the environment, or capabilities to operate in complex situations. Artificial Intelligence is extending the capability of perception and action of the agents and allows robots to operate in environments not suitable for robots just a few years ago. In most common scenarios the complexity of the environment requires to the robot to have different skills, the capability of different actions, and hence also a certain degree of reasoning and understanding of which action to take and when. A relevant example could be an urban road, with car, pedestrian, and signals.
Generating Safe Autonomous Decision-Making in ROS
The Robot Operating System (ROS) is a widely used framework for building robotic systems. It offers a wide variety of reusable packages and a pattern for new developments. It is up to developers how to combine these elements and integrate them with decision-making for autonomous behavior. The feature of such decision-making that is in general valued the most is safety assurance. In this research preview, we present a formal approach for generating safe autonomous decision-making in ROS. We first describe how to improve our existing static verification approach to verify multi-goal multi-agent decision-making. After that, we describe how to transition from the improved static verification approach to the proposed runtime verification approach. An initial implementation of this research proposal yields promising results.