Goto

Collaborating Authors

 Agents


c97e7a5153badb6576d8939469f58336-Supplemental.pdf

Neural Information Processing Systems

Our initial experiments (implementation, debugging, hyperparameter tuning, etc.) required about 5000CPUhoursofcompute. Due to these rules, it is recommended to group together in order to attack simultaneously. In Warehouse[4], QTRAN makes slightly faster progress than VAST(ฮท = 12). The results forWarehouse[16], Battle[80], and GaussianSqueeze[800] are shown in Figure 1. Figure 10: Visualizations of the generated sub-teams ofXMetaGrad with ฮท = 14 and XSpatial with k-means clustering using 10 centroids at different stages (early, middle, late) inBattle[80] after training. Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments.









f0eb6568ea114ba6e293f903c34d7488-Paper.pdf

Neural Information Processing Systems

Several works haveshown this vulnerability via adversarial attacks, butexisting approaches onimproving therobustness ofDRL under this setting have limited success and lack for theoretical principles. We show that naively applying existing techniques on improving robustness for classification tasks,likeadversarialtraining,areineffectiveformanyRLtasks.


c3e0c62ee91db8dc7382bde7419bb573-Supplemental.pdf

Neural Information Processing Systems

Theactiveagent trains (as a regular Double-DQN) up to the time of forking, at which point the passive agent is created asa'fork' (i.e.,with identical networkweights) oftheactiveagent.