Goto

Collaborating Authors

 Country




FunctionalEnsembleDistillation

Neural Information Processing Systems

One popular approach to alleviate this problem is using a Monte-Carlo estimation with an ensemble of models sampled from the posterior. However, this approach still comes at a significant computational cost, as one needs to store and run multiple models at testtime.





SupplementaryMaterial AutoSync: LearningtoSynchronizeforData-Parallel DistributedDeepLearning

Neural Information Processing Systems

Asδ is relatively slow compared with other time consumption, we treat it asaconstant that does not scale with variable size. The second process introduces network overhead (e.g.,latency) and network communication. For vi VCC, we model 5 mostly used collective primitives:AllReduce, ReduceScatter, AllGather, Broadcast and Reduce [12]. I1,I2,I3 are true whenAllReduce, ReduceScatter and AllGather, Broadcast and Reduce are activated, respectively. Note thattheLSTM walks through eachvi V0G,θ strictly following their original forward (backward) order in the computational graph, so as to inject this information intothemodeling. A global load balancer (clb) and group assigner (cam) assign their values using randomized and approximate solutions, illustrated in Algorithm 1 and Algorithm 2, respectively.



A Properties of coherent distortion risk measures

Neural Information Processing Systems

The properties of coherent risk measures also lead to a useful dual representation. Let ρ be a proper, real-valued coherent risk measure. See Shapiro et al. [42] for a general treatment of this result. Therefore, we have that the RAMU safe RL problem in (3) is equivalent to (6).B.3 Proof of Corollary 1 Fix ϵ > 0 and consider ( s, a) S A . Safety constraints and environment perturbations In all of our experiments, we consider the problem of optimizing a task objective while satisfying a safety constraint.


Risk-Averse Model Uncertainty for Distributionally Robust Safe Reinforcement Learning James Queeney

Neural Information Processing Systems

Many real-world domains require safe decision making in uncertain environments. In this work, we introduce a deep reinforcement learning framework for approaching this important problem. We consider a distribution over transition models, and apply a risk-averse perspective towards model uncertainty through the use of coherent distortion risk measures.