Goto

Collaborating Authors

 Country




ConservativeDualPolicyOptimizationforEfficient Model-Based ReinforcementLearning

Neural Information Processing Systems

Based ontheprinciple ofoptimism inthefaceofuncertainty(OFU) [56,49,10],OFU-RL achievestheglobal optimality by ensuring that the optimistically biased value is close to the real value in the long run. Based on Thompson Sampling [62], Posterior Sampling RL (PSRL) [57, 42, 43] explores by greedily optimizing the policy in an MDP which is sampled from the posterior distribution over MDPs.





Appendices

Neural Information Processing Systems

In this paper, we conduct experiments using six settings with Adam optimizer [18]. For the contrastive coefficientฮป (see Algorithm 1), the value is fixed at 1.0 for a fair comparison with [19, 8]. In all experiments, we use the temperaturet = 1.0. We stop training GANs with SNDCGAN, SNResGAN, and BigGAN architectures after 200k, 100k, and 80k generator updates, respectively. Experimental setup used for Table 3 in the main paper: FID values on CIFAR10 dataset are reported using the setting (E) with the batch size of 64.