Goto

Collaborating Authors

 Agents


An Accelerated Approach to Decentralized Reinforcement Learning of the Ball-Dribbling Behavior

AAAI Conferences

In the context of soccer robotics, ball dribbling is a complex behavior where a robot player attempts to maneuver the ball in a very controlled way, while moving towards a desired target. To learn when and how to modify the robotโ€™s velocity vector is a complex problem, hardly solvable in an effective way with methods based on identification of the system dynamics and/or kinematics and mathematical models. We propose a decentralized reinforcement learning strategy, where each component of the omnidirectional biped walk (𝑣𝑥,𝑣𝑦,𝑣𝜃) is learned in parallel with single-agents working in a multiagent task. Moreover, we propose an approach to accelerate the decentralized learning based on knowledge transfer from simple linear controllers. Obtained results are successful; with less human effort, and less required designer knowledge, the decentralized reinforcement learning scheme shows better performances than the current dribbling engine used by UChile Robotics Team in the SPL robot soccer competitions. The proposed decentralized rein- forcement learning scheme achieves asymptotic performance after 1500 episodes and can be accelerated up to 70% by using our approach to share actions.


Effect of Bundle Method in Distributed Lagrangian Relaxation Protocol

AAAI Conferences

The Generalized Mutual Assignment Problem (GMAP) is a maximization problem in distributed environments, where multiple agents select goods under resource constraints. Distributed Lagrangian Relaxation Protocols (DisLRP) are peer-to-peer communication protocols for solving GMAP instances. In DisLRPs, agents seek a good quality upper bound on the optimal value by solving the Lagrangian dual problem, which is a convex minimization problem. Existing DisLRPs exploit a subgradient method to explore a better upper bound by updating the Lagrange multipliers (prices) of goods. While the computational complexity of the subgradient method is very low, it cannot detect tha fact that an upper bound converges to the minimum. Moreover, solution oscillation sometimes occurs, which is critical for its performance. In this paper, we present a new DisLRP with a Bundle Method and refer to it as Bundle DisLRP (BDisLRP). The bundle method, which is also called the stabilized cutting planes method, has recently attracted much attention as a way to solve Lagrangian dual problems in centralized environments. We show that this method can also work in distributed environments. We experimentally compared BDisLRP with Adaptive DisLRP (ADisLRP), which is a previous protocol that exploits the subgradient method, to demonstrate that BDisLRP converged faster with better quality upper bounds than ADisLRP.


A New Perspective of Trust Through Multi-Attribute Auctions

AAAI Conferences

Auction mechanisms are very well known methods to allocate tasks when several agents are involved. Particularly, multi-attribute auctions are a special mechanism that allows the consideration of task attributes other than prices, such as delivery time or energy consumptions. Incentive compatible mechanisms encourage agents to reveal the attributes which agents estimate truthful, however, these mechanisms by themselves cannot know if such estimations are reliable or not due to uncertainty. Under such circumstances, trust could complement incentive compatibility reducing the risk of losses by the auctioneer. The use of trust in auctions is a well-studied problem; however, most of the works in the literature focus on how to model trust rather on how trust is used in the mechanism. Thus, this paper proposes an easy and systematic way to include a multi-faceted model of trust into multi-attribute auctions. Conversely to other previous works where trust is only used in the winner determination problem, the presented approach uses trust both in deciding the winner of the auction and in the payment to the corresponding bidder. According to the results obtained from the experimentation, the use of trust following the methodology presented in this paper highly reduces the number of winner bids from unreliable bidders and, therefore, the number of tasks executed in worse conditions than the agreed. Complementary, this paper proposes a new trust adaptation method which consists of increasing or decreasing the trust value (depending on whether the task is executed properly or not) according to a simple mathematical function with asymptotes on 0 and 1. This model does not present the rigidity problem present in other models of the literature when it comes to agents that have inconstant performances.


Strategyproof Mechanisms for One-Dimensional Hybrid and Obnoxious Facility Location Models

AAAI Conferences

We consider a strategic variant of the facility location problem. We would like to locate a facility on a closed interval. There are n agents located on that interval, divided into two types: type 1 agents, who wish for the facility to be as far from them as possible, and type 2 agents, who wish for the facility to be as close to them as possible. Our goal is to maximize a form of aggregated social benefit: maxisumโ€“ the sum of the agentsโ€™ utilities, or the egalitarian objectiveโ€“ the minimal agent utility. The strategic aspect of the problem is that the agentsโ€™ locations are not known to us, but rather reported to us by the agentsโ€“ an agent might misreport his location in an attempt to move the facility away from or towards to his true location. We therefore require the facility-locating mechanism to be strategyproof, namely that reporting truthfully is a dominant strategy for each agent. As simply maximizing the social benefit is generally not strategyproof, our goal is to design strategyproof mechanisms with good approximation ratios. In this paper, we provide a best-possible 3approximate deterministic strategyproof mechanism, as well as a 23/13 approximate randomized strategyproof mechanism, both for the maxisum objective. We provide lower bounds of 3 and 3/2 on the approximation ratio attainable for maxisum, in the deterministic and randomized settings, respectively. For the egalitarian objective, we show that no bounded approximation ratio is attainable in the deterministic setting, and provide a lower bound of 3/2 for the randomized setting. To obtain our deterministic lower bounds, we characterize all deterministic strategyproof mechanisms when all agents are of type 1. Finally, while still restricting ourselves to agents of type 1 only, we consider a generalized model that allows an agent to control more than one location. In this generalized model, we provide best-possible 3and 3 approximate strategyproof 2 mechanisms for the maxisum objective in the deterministic and randomized settings, respectively.


Flexibility Meets Variability: A Multiagent Constraint Based Approach for Incorporating Renewables into the Power Grid

AAAI Conferences

This paper outlines a new approach to creating value from the Smart Grid by incorporating individual households into the response system that must be deployed to accommodate increasingly large sources of intermittent renewable power. We propose a framework that couples agent-based AI techniques with envelope methods. Envelope methods provide a unified mathematical framework to model intermittent renewable resources, conventional dispatchable resources, demand side response, and storage. The overall goal of our system is to develop a distributed autonomous agent architecture that is able to facilitate market transactions among load serving entities, residential consumers, conventional merchant power producers, and intermittent power producers.


E-HBA: Using Action Policies for Expert Advice and Agent Typification

AAAI Conferences

Past research has studied two approaches to utilise pre-defined policy sets in repeated interactions: as experts, to dictate our own actions, and as types, to characterise the behaviour of other agents. In this work, we bring these complementary views together in the form of a novel meta-algorithm, called Expert-HBA (E-HBA), which can be applied to any expert algorithm that considers the average (or total) payoff an expert has yielded in the past. E-HBA gradually mixes the past payoff with a predicted future payoff, which is computed using the type-based characterisation. We present results from a comprehensive set of repeated matrix games, comparing the performance of several well-known expert algorithms with and without the aid of E-HBA. Our results show that E-HBA has the potential to significantly improve the performance of expert algorithms.


Real-Time Optimal Selection of Multirobot Coalition Formation Algorithms Using Conceptual Clustering

AAAI Conferences

The presented framework is the The multirobot coalition formation problem seeks to intelligently first to leverage a conceptual clustering technique to partition partition a team of heterogeneous robots into any set of coalition formation algorithms in order to derive coalitions for a set of real-world tasks. Besides being N Pan optimal hierarchy classification tree, given any classification complete (Sandholm et al. 1999), the problem is also hard taxonomy. The results contribute to the state-ofthe-art to approximate (Service and Adams 2011a). Traditional approaches in multiagent systems by demonstrating the existence to solving the problem include a number of greedy of crucial patterns and intricate relationships among existing algorithms (Shehory and Kraus 1998; Vig and Adams coalition algorithms.


Two Algorithms for the Movements of Robotic Bodyguard Teams

AAAI Conferences

In this paper we consider a scenario where one or more robotic bodyguards are protecting an important individual (VIP) moving in a public space against harassment or harm from unarmed civilians. In this scenario, the main objective of the robots is to position themselves such that at any given moment they provide maximum physical cover for the VIP. The robots need to follow the VIP in its movement and take into account the movements of the civilians as well. The environment can also contain obstacles which present challenges to movement but also provide natural cover. We designed two algorithms for the movement of the bodyguard robots: Threat Vector Resolution (TVR) for a single robot and Quadrant Load Balancing (QLB) for teams of bodyguard robots. We evaluated the proposed approaches against rigid formations in a simulation study.


Trust, Influence and Reputation Management Based on Human Reasoning

AAAI Conferences

Understanding trust, influence and reputation and constructing computational models of these notions are two essential scientific challenges in computer science as well as social sciences. Although scientists in both disciplines have independently conducted research on these topics over the last couple of decades, there is a huge gap between two literatures. This paper therefore illustrates an interdisciplinary work-in-progress on trust, influence and reputation modeling based on human reasoning. Using a survey-based data collection approach, we would like to understand how humans gain/lose trust in their daily life interactions and how behavior/attitudes of humans can be influenced or shaped in various social encounters. The data will be then transformed into mathematical models to be used in technological or software systems.


Agents Vote for the Environment: Designing Energy-Efficient Architecture

AAAI Conferences

Saving energy is a major concern. Hence, it is fundamental to design and construct buildings that are energy-efficient. It is known that the early stage of architectural design has a significant impact on this matter. However, it is complex to create designs that are optimally energy efficient, and at the same time balance other essential criterias such as economics, space, and safety. One state-of-the art approach is to create parametric designs, and use a genetic algorithm to optimize across different objectives. We further improve this method, by aggregating the solutions of multiple agents. We evaluate diverse teams, composed by different agents; and uniform teams, composed by multiple copies of a single agent. We test our approach across three design cases of increasing complexity, and show that the diverse team provides a significantly larger percentage of optimal solutions than single agents.