Goto

Collaborating Authors

 Agents



FindingRegionsofHeterogeneityinDecision-Making viaExpectedConditionalCovariance

Neural Information Processing Systems

Individuals often make different decisions when faced with the same context, due to personal preferences and background. For instance, judges may vary in their leniency towards certain drug-related offenses, and doctors may vary in their preference for how to start treatment for certain types of patients.


Multi-AgentReinforcementLearningis ASequenceModelingProblem

Neural Information Processing Systems

Recently, such difficulty in multi-agent learning has been eased owing to the introduction ofcentralized training for decentralized execution(CTDE) [11, 45], which allows agents to access the global information andopponents' actions during thetraining phase.






Multi-agentactiveperceptionwithpredictionrewards

Neural Information Processing Systems

Active perception,collecting observations to reduce uncertainty about ahidden variable, isone of the fundamental capabilities of an intelligent agent [2]. In multi-agent active perceptiona team of autonomous agents cooperatively gathers observations to infer the value of a hidden variable.



Agent 1 Agent 2 River Tiles (a) The initial setup with two agents and two river

Neural Information Processing Systems

Agent 1's action is resolved first. Figure 8: An example of Agent 1 using the "clean" action while facing East. The "main" beam extends directly in front of the agent, while two auxiliary A beam stops when it hits a dirty river tile. The Sequential Social Dilemma Games, introduced in Leibo et al. [2017], are a kind of MARL All of these have open source implementations in [Vinitsky et al., 2019]. The cleaning beam is shown in Figure 8a.