Agents
Towards Graph Representation Learning in Emergent Communication
Słowik, Agnieszka, Gupta, Abhinav, Hamilton, William L., Jamnik, Mateja, Holden, Sean B.
Recent findings in neuroscience suggest that the human brain represents information in a geometric structure (for instance, through conceptual spaces). In order to communicate, we flatten the complex representation of entities and their attributes into a single word or a sentence. In this paper we use graph convolutional networks to support the evolution of language and cooperation in multi-agent systems. Motivated by an image-based referential game, we propose a graph referential game with varying degrees of complexity, and we provide strong baseline models that exhibit desirable properties in terms of language emergence and cooperation. We show that the emerged communication protocol is robust, that the agents uncover the true factors of variation in the game, and that they learn to generalize beyond the samples encountered during training.
New Research Hints at How Your Smiles Could One Day Teach Artificial Intelligence
We live in an era when humans are busy training a new intelligence on this planet. Every once in a while, researchers come up with a novel way to speed up that teaching process. That's what happened at Microsoft Research where computer scientists recently developed a new approach to using human emotion to train machines how to learn.[i] The research used virtual agents to facilitate learning various tasks in a simulated environment. What is most significant about this research is that it trained those agents by exposing them to the smiles of human subjects as they interacted with the system.
I Feel I Feel You: A Theory of Mind Experiment in Games
Melhart, David, Yannakakis, Georgios N., Liapis, Antonios
In this study into the player's emotional theory of mind of gameplaying agents, we investigate how an agent's behaviour and the player's own performance and emotions shape the recognition of a frustrated behaviour. We focus on the perception of frustration as it is a prevalent affective experience in human-computer interaction. We present a testbed game tailored towards this end, in which a player competes against an agent with a frustration model based on theory. We collect gameplay data, an annotated ground truth about the player's appraisal of the agent's frustration, and apply face recognition to estimate the player's emotional state. We examine the collected data through correlation analysis and predictive machine learning models, and find that the player's observable emotions are not correlated highly with the perceived frustration of the agent. This suggests that our subject's theory of mind is a cognitive process based on the gameplay context. Our predictive models---using ranking support vector machines---corroborate these results, yielding moderately accurate predictors of players' theory of mind.
Compositional properties of emergent languages in deep learning
Keresztury, Bence, Bruni, Elia
Usually two or more agents play a cooperative game where agents' goals are common but they access different information. In order to solve the given task agents have to share useful information with each other through a discrete bottleneck called the communication channel. The discrete symbols in the message do not have any a priori meaning but agents learn to cooperate by attributing meaning to the messages; a language protocol emerges as a byproduct of the training process. This emergent language serves only one goal: to complete the task successfully. One of the promises of this approach is to provide meaningful insights into the early stages of human language emergence as a result of cooperation.
On Solving Cooperative MARL Problems with a Few Good Experiences
Kumar, Rajiv Ranjan, Varakantham, Pradeep
Cooperative Multi-agent Reinforcement Learning (MARL) is crucial for cooperative decentralized decision learning in many domains such as search and rescue, drone surveillance, package delivery and fire fighting problems. In these domains, a key challenge is learning with a few good experiences, i.e., positive reinforcements are obtained only in a few situations (e.g., on extinguishing a fire or tracking a crime or delivering a package) and in most other situations there is zero or negative reinforcement. Learning decisions with a few good experiences is extremely challenging in cooperative MARL problems due to three reasons. First, compared to the single agent case, exploration is harder as multiple agents have to be coordinated to receive a good experience. Second, environment is not stationary as all the agents are learning at the same time (and hence change policies). Third, scale of problem increases significantly with every additional agent. Relevant existing work is extensive and has focussed on dealing with a few good experiences in single-agent RL problems or on scalable approaches for handling non-stationarity in MARL problems. Unfortunately, neither of these approaches (or their extensions) are able to address the problem of sparse good experiences effectively. Therefore, we provide a novel fictitious self imitation approach that is able to simultaneously handle non-stationarity and sparse good experiences in a scalable manner. Finally, we provide a thorough comparison (experimental or descriptive) against relevant cooperative MARL algorithms to demonstrate the utility of our approach.
Numerical Abstract Persuasion Argumentation for Expressing Concurrent Multi-Agent Negotiations
A negotiation process by 2 agents e1 and e2 can be interleaved by another negotiation process between, say, e1 and e3. The interleaving may alter the resource allocation assumed at the inception of the first negotiation process. Existing proposals for argumentation-based negotiations have focused primarily on two-agent bilateral negotiations, but scarcely on the concurrency of multi-agent negotiations. To fill the gap, we present a novel argumentation theory, basing its development on abstract persuasion argumentation (which is an abstract argumentation formalism with a dynamic relation). Incorporating into it numerical information and a mechanism of handshakes among members of the dynamic relation, we show that the extended theory adapts well to concurrent multi-agent negotiations over scarce resources.
Proxy Tasks and Subjective Measures Can Be Misleading in Evaluating Explainable AI Systems
Buçinca, Zana, Lin, Phoebe, Gajos, Krzysztof Z., Glassman, Elena L.
Explainable artificially intelligent (XAI) systems form part of sociotechnical systems, e.g., human+AI teams tasked with making decisions. Yet, current XAI systems are rarely evaluated by measuring the performance of human+AI teams on actual decision-making tasks. We conducted two online experiments and one in-person think-aloud study to evaluate two currently common techniques for evaluating XAI systems: (1) using proxy, artificial tasks such as how well humans predict the AI's decision from the given explanations, and (2) using subjective measures of trust and preference as predictors of actual performance. The results of our experiments demonstrate that evaluations with proxy tasks did not predict the results of the evaluations with the actual decision-making tasks. Further, the subjective measures on evaluations with actual decision-making tasks did not predict the objective performance on those same tasks. Our results suggest that by employing misleading evaluation methods, our field may be inadvertently slowing its progress toward developing human+AI teams that can reliably perform better than humans or AIs alone.
Subjective Knowledge and Reasoning about Agents in Multi-Agent Systems
Singh, Shikha, Khemani, Deepak
Though a lot of work in multi-agent systems is focused on reasoning about knowledge and beliefs of artificial agents, an explicit representation and reasoning about the presence/absence of agents, especially in the scenarios where agents may be unaware of other agents joining in or going offline in a multi-agent system, leading to partial knowledge/asymmetric knowledge of the agents is mostly overlooked by the MAS community. Such scenarios lay the foundations of cases where an agent can influence other agents' mental states by (mis)informing them about the presence/absence of collaborators or adversaries. In this paper, we investigate how Kripke structure-based epistemic models can be extended to express the above notion based on an agent's subjective knowledge and we discuss the challenges that come along.
Understanding the key differences between chatbots and virtual agents
The differences between a virtual agent and a chatbot are actually bigger than you might think. To help distinguish between the two technologies, it's helpful to draw a parallel with another popular technology--the smartphone. When it comes to understanding the difference between chatbots and virtual agents, there are parallels to the evolution of a technology that has evolved significantly and is not referred to differently than it used to be--the smartphone. Fewer and fewer people regularly use the words'telephone' and'smartphone' interchangeably anymore, primarily because they are technically and functionally referring to two very different devices. Both a phone and a smartphone can be used to make calls, but that's where the similarities stop.
On Algorithmic Decision Procedures in Emergency Response Systems in Smart and Connected Communities
Pettet, Geoffrey, Mukhopadhyay, Ayan, Kochenderfer, Mykel, Vorobeychik, Yevgeniy, Dubey, Abhishek
Emergency Response Management (ERM) is a critical problem faced by communities across the globe. Despite its importance, it is common for ERM systems to follow myopic and straight-forward decision policies in the real world. Principled approaches to aid decision-making under uncertainty have been explored in this context but have failed to be accepted into real systems. We identify a key issue impeding their adoption - algorithmic approaches to emergency response focus on reactive, post-incident dispatching actions, i.e. optimally dispatching a responder after incidents occur. However, the critical nature of emergency response dictates that when an incident occurs, first responders always dispatch the closest available responder to the incident. We argue that the crucial period of planning for ERM systems is not post-incident, but between incidents. However, this is not a trivial planning problem - a major challenge with dynamically balancing the spatial distribution of responders is the complexity of the problem. An orthogonal problem in ERM systems is to plan under limited communication, which is particularly important in disaster scenarios that affect communication networks. We address both the problems by proposing two partially decentralized multi-agent planning algorithms that utilize heuristics and the structure of the dispatch problem. We evaluate our proposed approach using real-world data, and find that in several contexts, dynamic re-balancing the spatial distribution of emergency responders reduces both the average response time as well as its variance.