Agents
Multi-agent Hierarchical Reinforcement Learning with Dynamic Termination
Han, Dongge, Boehmer, Wendelin, Wooldridge, Michael, Rogers, Alex
In a multi-agent system, an agent's optimal policy will typically depend on the policies chosen by others. Therefore, a key issue in multi-agent systems research is that of predicting the behaviours of others, and responding promptly to changes in such behaviours. One obvious possibility is for each agent to broadcast their current intention, for example, the currently executed option in a hierarchical reinforcement learning framework. However, this approach results in inflexibility of agents if options have an extended duration and are dynamic. While adjusting the executed option at each step improves flexibility from a single-agent perspective, frequent changes in options can induce inconsistency between an agent's actual behaviour and its broadcast intention. In order to balance flexibility and predictability, we propose a dynamic termination Bellman equation that allows the agents to flexibly terminate their options. We evaluate our model empirically on a set of multi-agent pursuit and taxi tasks, and show that our agents learn to adapt flexibly across scenarios that require different termination behaviours.
Redistribution Mechanism Design on Networks
Zhang, Wen, Zhao, Dengji, Chen, Hanyu
Redistribution mechanisms have been proposed for more efficient resource allocation but not for profit. We consider redistribution mechanism design for the first time in a setting where participants are connected and the resource owner is only aware of her neighbours. In this setting, to make the resource allocation more efficient, the resource owner has to inform the others who are not her neighbours, but her neighbours do not want more participants to compete with them. Hence, the goal is to design a redistribution mechanism such that participants are incentivized to invite more participants and the resource owner does not earn or lose much money from the allocation. We first show that existing redistribution mechanisms cannot be directly applied in the network setting to achieve the goal. Then we propose a novel network-based redistribution mechanism such that all participants in the network are invited, the allocation is more efficient and the resource owner has no deficit. Introduction The problem of resource allocation has recently caught the public imagination, where the resource owner has to decide the allocation of the item among a group of self-interested agents. Since the valuation differs from agents, it is a natural objective for the owner to pursue the efficiency of the allocation, i.e., allocating the item to the agent with the highest valuation. In many scenarios, the owner does not really aim at making profits but hopes the wealth maintained among the agents. For example, the government wants to build a library in a community that values it most; a charity distributes a donation to the recipient who needs it most; a hospital allocates doctors to rural areas where doctors are highly demanded. To find the agent with the highest valuation, one common alternative is to hold an auction (Krishna 2009) under some protocols such as the well-known Vickrey-Clarke- Groves (VCG) mechanism (Vickrey 1961; Clarke 1971; Groves 1973). However, the payments under VCG will all be delivered to the auctioneer, which againsts our nonprofit purpose.
Autonomous Industrial Management via Reinforcement Learning: Self-Learning Agents for Decision-Making -- A Review
Leal, Leonardo A. Espinosa, Westerlund, Magnus, Chapman, Anthony
Industry has always been in the pursuit of becoming more economically efficient and the current focus has been to reduce human labour using modern technologies. Even with cutting edge technologies, which range from packaging robots to AI for fault detection, there is still some ambiguity on the aims of some new systems, namely, whether they are automated or autonomous. In this paper we indicate the distinctions between automated and autonomous system as well as review the current literature and identify the core challenges for creating learning mechanisms of autonomous agents. We discuss using different types of extended realities, such as digital twins, to train reinforcement learning agents to learn specific tasks through generalization. Once generalization is achieved, we discuss how these can be used to develop self-learning agents. We then introduce self-play scenarios and how they can be used to teach self-learning agents through a supportive environment which focuses on how the agents can adapt to different real-world environments.
Leverage AI to Create Autonomous Policies that Adapts without Human Intervention
Policies are the foundation for any successful organization. Policies are the rules, or laws, of an organization. Heck, one could argue that an organization's culture is better defined by its policies than it is by the character of its leadership team. Unfortunately, the management, creation and execution of policies haven't changed much since the days of "time-and-motion studies". In many cases, policies are nothing more than a static list of what-if rules that govern what workers are to do in well-defined situations.
A Structured Prediction Approach for Generalization in Cooperative Multi-Agent Reinforcement Learning
Carion, Nicolas, Synnaeve, Gabriel, Lazaric, Alessandro, Usunier, Nicolas
Effective coordination is crucial to solve multi-agent collaborative (MAC) problems. While centralized reinforcement learning methods can optimally solve small MAC instances, they do not scale to large problems and they fail to generalize to scenarios different from those seen during training. In this paper, we consider MAC problems with some intrinsic notion of locality (e.g., geographic proximity) such that interactions between agents and tasks are locally limited. By leveraging this property, we introduce a novel structured prediction approach to assign agents to tasks. At each step, the assignment is obtained by solving a centralized optimization problem (the inference procedure) whose objective function is parameterized by a learned scoring model. We propose different combinations of inference procedures and scoring models able to represent coordination patterns of increasing complexity. The resulting assignment policy can be efficiently learned on small problem instances and readily reused in problems with more agents and tasks (i.e., zero-shot generalization). We report experimental results on a toy search and rescue problem and on several target selection scenarios in StarCraft: Brood War, in which our model significantly outperforms strong rule-based baselines on instances with 5 times more agents and tasks than those seen during training.
How to Implement a Ticket Triaging System with AI
Customer queries are the bane of most customer support teams, not because they don't like dealing with them, but because they don't have a proper process in place that lets them handle excessive ticket volumes easily and effectively. When a support ticket drops into a queue, or an agent receives an email with a customer issue, the ticket or email might pass through three different agents before finally landing in the correct hands to deal with the issue โ leading to bottlenecks and bad customer experiences. Bugs, forgotten passwords, system errors, integration queriesโฆ There are so many different issues that agents have to deal with, so that the customer remains happy and the company retains them. And while customer support endeavors to respond to queries as quickly as possible, it's difficult when faced with huge volumes of tickets. On top of that, more and more customers expect immediate responses โ 64% of consumers and 80% of business buyers said they expect companies to respond to and interact with them in real time. Deciding how to tackle customer requests, which to tackle first, and making sure tickets are sent to the right person โ or team that's best equipped to deal with the query โ are processes that need to run as smoothly as possible, so that organizations can score high on the customer satisfaction scale.
Ten Tips For Deploying Enterprise Virtual Agents
We use artificial intelligence (AI) every day without knowing it. Alexa's speech recognition, Netflix's movie recommendations and Gmail's type-ahead suggestions are all examples of data you share being fed into deep learning AI models to improve your experience. At work, the same technology is increasingly used to deliver smart services -- rooms that book themselves, thermostats that don't cool empty floors and outages that are restored before anyone knows they exist. Achieving Amazon-quality AI with "small data" from your organization is a more complex technical problem. To get enterprise AI right at scale requires thinking differently about how applications and services are deployed and managed.
Blameworthiness in Security Games
Security games are an example of a successful real-world application of game theory. The paper defines blameworthiness of the defender and the attacker in security games using the principle of alternative possibilities and provides a sound and complete logical system for reasoning about blameworthiness in such games. Introduction In this paper we study the properties of blameworthiness in security games (von Stackelberg 1934). Security games are used for canine airport patrol (Pita et al. 2008; Jain et al. 2010), airport passenger screening (Brown et al. 2016), protecting endangered animals and fish stocks (Fang, Stone, and Tambe 2015), U.S. Coast Guard port patrol (Sinha et al. 2018; An, Tambe, and Sinha 2016), and randomized deployment of U.S. air marshals (Sinha et al. 2018). Defender \Attacker Terminal 1 Terminal 2 Terminal 1 20 120 Terminal 2 200 16 Figure 1: Expected Human Losses in Security Game G 1. As an example, consider a security game G 1 in which a defender is trying to protect two terminals in an airport from an attacker. Due to limited resources, the defender can patrol only one terminal at a given time. If the defender chooses to patrol Terminal 1 and the attacker chooses to attack Terminal 2, then the human losses at Terminal 2 are estimated at 120, see Figure 1. However, if the defender chooses to patrol Terminal 2 while the attacker still chooses to attack Terminal 2, then the expected number of the human losses at Terminal 2 is only 16, see Figure 1. Generally speaking, the goal of the defender is to minimize human losses, while the goal of the attacker is to maximize them. However, the utility functions in security games usually take into account not only the human losses, but also the cost to protect and to attack the target to the defender and the attacker respectively.
I visualised how algorithms 'see' urban environments and build detailed profiles of citizens
Algorithms, software and smart technologies have a growing presence in cities around the world. Artificial intelligence (AI), agent-based modelling, the internet of things and machine learning can be found practically everywhere now โ from lampposts to garbage bins, traffic lights and cars. Not only that, these technologies are also influencing how cities are planned, guiding big decisions about new buildings, transport and infrastructure projects. City-dwellers tend to accept the presence of these technologies passively โ if they notice it at all. Yet this acceptance is punctuated by intermittent panic over privacy โ take, for example, Transport for London's latest plans to track passenger journeys across the transport network using wifi, which drew criticism from privacy experts.
Do we trust artificial intelligence agents to mediate conflict? Not entirely: New study says we'll listen to virtual agents except when goings get tough
Researchers from USC and the University of Denver created a simulation in which a three-person team was supported by a virtual agent avatar on screen in a mission that was designed to ensure failure and elicit conflict. The study was designed to look at virtual agents as potential mediators to improve team collaboration during conflict mediation. But in the heat of the moment, will we listen to virtual agents? While some of researchers (Gale Lucas and Jonathan Gratch of the USC Viterbi School Engineering and the USC Institute for Creative Technologies who contributed to this study), had previously found that one-on-one human interactions with a virtual agent therapist yielded more confessions, in this study "Conflict Mediation in Human-Machine Teaming: Using a Virtual Agent to Support Mission Planning and Debriefing," team members were less likely to engage with a male virtual agent named "Chris" when conflict arose. Participating members of the team did not physically accost the device (as we have seen humans attack robots in viral social media posts), but rather were less engaged and less likely to listen to the virtual agent's input once failure ensued and conflict arose among team members. The study was conducted in a military academy environment in which 27 scenarios were engineered to test how the team that included a virtual agent would react to failure and the ensuring conflict.