Goto

Collaborating Authors

 Country


a3621ee907def47c1b952ade25c67698-Paper-Conference.pdf

Neural Information Processing Systems

This paper explores the potential of building scalable techniques to facilitate autonomous cooperation among communicative agents, and provides insight into their "cognitive" processes. To address the challenges of achieving autonomous cooperation, we propose a novel communicative agent framework named roleplaying .








A Proof Proof of Proposition 4.2 Proposition 4.2 The performance gap of evaluating policy profile (ฯ€, ยต) and (ฯ€, ฯ€

Neural Information Processing Systems

Proof of Theorem 4.7 We first prove a Lemma. Theorem A.2. (Theorem 1 in [36]) Let ฯต = max Theorem 4.7 In a two-player game, suppose that According to Theorem A.2, we have J ( ฯ€, ยต) J ( ฯ€, ฮฑ) E CQL [20] puts regularization on the learning of Q function to penalize out-of-distribution actions. The CSP algorithm is illustrated in Algorithm 1. The proxy model is trained adversarially against our agent, therefore, we set the proxy's reward function to be the negative of our agent's reward. We show experiment details of the Maze example in this section.



FouRA: Fourier Low Rank Adaptation

Neural Information Processing Systems

While Low-Rank Adaptation (LoRA) has proven beneficial for efficiently fine-tuning large models, LoRA fine-tuned text-to-image diffusion models lack diversity in the generated images, as the model tends to copy data from the observed training samples.