Agents
Does Spending More Always Ensure Higher Cooperation? An Analysis of Institutional Incentives on Heterogeneous Networks
Cimpeanu, Theodor, Santos, Francisco C, Han, The Anh
Humans have developed considerable machinery used at scale to create policies and to distribute incentives, yet we are forever seeking ways in which to improve upon these, our institutions. Especially when funding is limited, it is imperative to optimise spending without sacrificing positive outcomes, a challenge which has often been approached within several areas of social, life and engineering sciences. These studies often neglect the availability of information, cost restraints, or the underlying complex network structures, which define real-world populations. Here, we have extended these models, including the aforementioned concerns, but also tested the robustness of their findings to stochastic social learning paradigms. Akin to real-world decisions on how best to distribute endowments, we study several incentive schemes, which consider information about the overall population, local neighbourhoods, or the level of influence which a cooperative node has in the network, selectively rewarding cooperative behaviour if certain criteria are met. Following a transition towards a more realistic network setting and stochastic behavioural update rule, we found that carelessly promoting cooperators can often lead to their downfall in socially diverse settings. These emergent cyclic patterns not only damage cooperation, but also decimate the budgets of external investors. Our findings highlight the complexity of designing effective and cogent investment policies in socially diverse populations.
Optimization Algorithms in Smart Grids: A Systematic Literature Review
Aslam, Sidra, Altaweel, Ala, Nassif, Ali Bou
Abstract--Electrical smart grids are units that supply electricity from power plants to the users to yield reduced costs, power failures/loss, and maximized energy management. Smart grids (SGs) are well-known devices due to their exceptional benefits such as bi-directional communication, stability, detection of power failures, and inter-connectivity with appliances for monitoring purposes. Hence, the importance of SGs as a research field is increasing with every passing year. This paper focuses on novel features and applications of smart grids in domestic and industrial sectors. Specifically, we focused on Genetic algorithm, Particle Swarm Optimization, and Grey Wolf Optimization to study the efforts made up till date for maximized energy management and cost minimization in SGs. Many counter Smart grids refers to an electric grid that delivers the attack solutions such as secure data collectors, broadcast authentication, electricity from utility (power generator sources/company) to and secure DoS-resistant broadcast authentication the users (residential/industrial). A simple smart grid connection protocols have been studied to secure the data collection and is shown in Figure 1, with bi-directional communication coping the demands of users in efficient ways [9], [10]. The process of electricity other challenges are faced by both utility and users (energy delivery is capable of monitoring, modeling, controlling, data supply and energy demand) such as energy management, filtering, and data processing with help of number of intelligent cost efficiency, reducing power losses, and reducing pollutant features such as Artificial Intelligence (AI) or Computational emissions [11], [12]. The aforementioned challenges can be Intelligence (CI) as shown in Figure 2. SGs allow users to addressed using optimization techniques in SGs to maximize schedule the appliances depending upon pricing hours and the profit (for both users and utility) by managing electricity its demand that helps in saving energy, increasing reliability, distribution and reducing emissions. Furthermore, SGs support Optimization in SGs is employed to find the conditions with bidirectional power line communications such as Home Area maximum benefits while (at the same time) minimizing the Network (HAN) or Wide Area Network (WAN), and wireless electricity wastage and cost [13]. Hence, optimization problem communications such as ZigBee, 6LowPAN, Z-wave, IoT in SGs is defined as a scenario (i.e., an objective function) that networks, etc. [3]-[6]. For future work, we aim to expand our research for other optimization algorithms (i.e., ABC, ACO). Our contributions in this paper are: fluenced by a set of variables and/or constraints.
Cooperative Concurrent Games
Gutierrez, Julian, Kowara, Szymon, Kraus, Sarit, Steeples, Thomas, Wooldridge, Michael
In rational verification, the aim is to verify which temporal logic properties will obtain in a multi-agent system, under the assumption that agents ("players") in the system choose strategies for acting that form a game theoretic equilibrium. Preferences are typically defined by assuming that agents act in pursuit of individual goals, specified as temporal logic formulae. To date, rational verification has been studied using non-cooperative solution concepts - Nash equilibrium and refinements thereof. Such non-cooperative solution concepts assume that there is no possibility of agents forming binding agreements to cooperate, and as such they are restricted in their applicability. In this article, we extend rational verification to cooperative solution concepts, as studied in the field of cooperative game theory. We focus on the core, as this is the most fundamental (and most widely studied) cooperative solution concept. We begin by presenting a variant of the core that seems well-suited to the concurrent game setting, and we show that this version of the core can be characterised using ATL*. We then study the computational complexity of key decision problems associated with the core, which range from problems in PSPACE to problems in 3EXPTIME. We also investigate conditions that are sufficient to ensure that the core is non-empty, and explore when it is invariant under bisimilarity. We then introduce and study a number of variants of the main definition of the core, leading to the issue of credible deviations, and to stronger notions of collective stable behaviour. Finally, we study cooperative rational verification using an alternative model of preferences, in which players seek to maximise the mean-payoff they obtain over an infinite play in games where quantitative information is allowed.
Opponent-aware Role-based Learning in Team Competitive Markov Games
Koley, Paramita, Maiti, Aurghya, Ganguly, Niloy, Bhattacharya, Sourangshu
Team competition in multi-agent Markov games is an increasingly important setting for multi-agent reinforcement learning, due to its general applicability in modeling many real-life situations. Multi-agent actor-critic methods are the most suitable class of techniques for learning optimal policies in the team competition setting, due to their flexibility in learning agent-specific critic functions, which can also learn from other agents. In many real-world team competitive scenarios, the roles of the agents naturally emerge, in order to aid in coordination and collaboration within members of the teams. However, existing methods for learning emergent roles rely heavily on the Q-learning setup which does not allow learning of agent-specific Q-functions. In this paper, we propose RAC, a novel technique for learning the emergent roles of agents within a team that are diverse and dynamic. In the proposed method, agents also benefit from predicting the roles of the agents in the opponent team. RAC uses the actor-critic framework with role encoder and opponent role predictors for learning an optimal policy. Experimentation using 2 games demonstrates that the policies learned by RAC achieve higher rewards than those learned using state-of-the-art baselines. Moreover, experiments suggest that the agents in a team learn diverse and opponent-aware policies.
FeSAC: Federated Learning-Based Soft Actor-Critic Traffic Offloading in Space-Air-Ground Integrated Network
Tang, Fengxiao, Yang, Yilin, Yao, Xin, Zhao, Ming, Kato, Nei
With the increase of intelligent devices leading to increasing demand for traffic, traffic offloading has become a challenging problem. The space-air-ground integrated network (SAGIN) is a superior network architecture to solve this problem. The existing research on SAGIN traffic offloading only considers the single-layer satellite network in the space network. To further expand the resource pool of traffic offloading in SAGIN, we extend the single-layer satellite network into a double-layer satellite network composed of low-orbit satellites (LEO) and high-orbit satellites (GEO). And re-model a four-layer SAGIN architecture consisting of the ground network, the air network, LEO and GEO. Furthermore, we propose a novel Federated Soft Actor-Critic (FeSAC) traffic offloading method with positive environmental exploration to accommodate this dynamic and complex four-layer SAGIN architecture. The FeSAC method uses federated learning to train traffic offloading nodes and then aggregate the training results to obtain the best traffic offloading strategy. The simulation results show that under the four-layer SAGIN, our proposed method can better adapt to the network environment changes by nodes mobility and is better than the existing traffic offloading methods in throughput, packet loss, and transmission delay.
Fairness and Sequential Decision Making: Limits, Lessons, and Opportunities
Nashed, Samer B., Svegliato, Justin, Blodgett, Su Lin
As automated decision making and decision assistance systems become common in everyday life, research on the prevention or mitigation of potential harms that arise from decisions made by these systems has proliferated. However, various research communities have independently conceptualized these harms, envisioned potential applications, and proposed interventions. The result is a somewhat fractured landscape of literature focused generally on ensuring decision-making algorithms "do the right thing". In this paper, we compare and discuss work across two major subsets of this literature: algorithmic fairness, which focuses primarily on predictive systems, and ethical decision making, which focuses primarily on sequential decision making and planning. We explore how each of these settings has articulated its normative concerns, the viability of different techniques for these different settings, and how ideas from each setting may have utility for the other.
Coordinated Multi-Robot Trajectory Tracking Control over Sampled Communication
Rossi, Enrica, Tognon, Marco, Ballotta, Luca, Carli, Ruggero, Cortรฉs, Juan, Franchi, Antonio, Schenato, Luca
In this paper, we propose an inverse-kinematics controller for a class of multi-robot systems in the scenario of sampled communication. The goal is to make a group of robots perform trajectory tracking in a coordinated way when the sampling time of communications is much larger than the sampling time of low-level controllers, disrupting theoretical convergence guarantees of standard control design in continuous time. Given a desired trajectory in configuration space which is precomputed offline, the proposed controller receives configuration measurements, possibly via wireless, to re-compute velocity references for the robots, which are tracked by a low-level controller. We propose joint design of a sampled proportional feedback plus a novel continuous-time feedforward that linearizes the dynamics around the reference trajectory: this method is amenable to distributed communication implementation where only one broadcast transmission is needed per sample. Also, we provide closed-form expressions for instability and stability regions and convergence rate in terms of proportional gain $k$ and sampling period $T$. We test the proposed control strategy via numerical simulations in the scenario of cooperative aerial manipulation of a cable-suspended load using a realistic simulator (Fly-Crane). Finally, we compare our proposed controller with centralized approaches that adapt the feedback gain online through smart heuristics, and show that it achieves comparable performance.
Guaranteed Encapsulation of Targets with Unknown Motion by a Minimalist Robotic Swarm
Sinhmar, Himani, Kress-Gazit, Hadas
We present a decentralized control algorithm for a robotic swarm given the task of encapsulating static and moving targets in a bounded unknown environment. We consider minimalist robots without memory, explicit communication, or localization information. The state-of-the-art approaches generally assume that the robots in the swarm are able to detect the relative position of neighboring robots and targets in order to provide convergence guarantees. In this work, we propose a novel control law for the guaranteed encapsulation of static and moving targets while avoiding all collisions, when the robots do not know the exact relative location of any robot or target in the environment. We make use of the Lyapunov stability theory to prove the convergence of our control algorithm and provide bounds on the ratio between the target and robot speeds. Furthermore, our proposed approach is able to provide stochastic guarantees under the bounds that we determine on task parameters for scenarios where a target moves faster than a robot. Finally, we present an analysis of how the emergent behavior changes with different parameters of the task and noisy sensor readings.
Multi-Target Landmark Detection with Incomplete Images via Reinforcement Learning and Shape Prior
Wan, Kaiwen, Li, Lei, Jia, Dengqiang, Gao, Shangqi, Qian, Wei, Wu, Yingzhi, Lin, Huandong, Mu, Xiongzheng, Gao, Xin, Wang, Sijia, Wu, Fuping, Zhuang, Xiahai
Medical images are generally acquired with limited field-of-view (FOV), which could lead to incomplete regions of interest (ROI), and thus impose a great challenge on medical image analysis. This is particularly evident for the learning-based multi-target landmark detection, where algorithms could be misleading to learn primarily the variation of background due to the varying FOV, failing the detection of targets. Based on learning a navigation policy, instead of predicting targets directly, reinforcement learning (RL)-based methods have the potential totackle this challenge in an efficient manner. Inspired by this, in this work we propose a multi-agent RL framework for simultaneous multi-target landmark detection. This framework is aimed to learn from incomplete or (and) complete images to form an implicit knowledge of global structure, which is consolidated during the training stage for the detection of targets from either complete or incomplete test images. To further explicitly exploit the global structural information from incomplete images, we propose to embed a shape model into the RL process. With this prior knowledge, the proposed RL model can not only localize dozens of targetssimultaneously, but also work effectively and robustly in the presence of incomplete images. We validated the applicability and efficacy of the proposed method on various multi-target detection tasks with incomplete images from practical clinics, using body dual-energy X-ray absorptiometry (DXA), cardiac MRI and head CT datasets. Results showed that our method could predict whole set of landmarks with incomplete training images up to 80% missing proportion (average distance error 2.29 cm on body DXA), and could detect unseen landmarks in regions with missing image information outside FOV of target images (average distance error 6.84 mm on 3D half-head CT).
Heterogeneous Beliefs and Multi-Population Learning in Network Games
Hu, Shuyue, Soh, Harold, Piliouras, Georgios
The effect of population heterogeneity in multi-agent learning is practically relevant but remains far from being well-understood. Motivated by this, we introduce a model of multi-population learning that allows for heterogeneous beliefs within each population and where agents respond to their beliefs via smooth fictitious play (SFP).We show that the system state -- a probability distribution over beliefs -- evolves according to a system of partial differential equations akin to the continuity equations that commonly desccribe transport phenomena in physical systems. We establish the convergence of SFP to Quantal Response Equilibria in different classes of games capturing both network competition as well as network coordination. We also prove that the beliefs will eventually homogenize in all network games. Although the initial belief heterogeneity disappears in the limit, we show that it plays a crucial role for equilibrium selection in the case of coordination games as it helps select highly desirable equilibria. Contrary, in the case of network competition, the resulting limit behavior is independent of the initialization of beliefs, even when the underlying game has many distinct Nash equilibria.