Goto

Collaborating Authors

 Country


HybridRegretBoundsforCombinatorial Semi-BanditsandAdversarialLinearBandits

Neural Information Processing Systems

Theformer means that the algorithm will work nearly optimally in all environments in an adversarial setting, a stochastic setting, or a stochastic setting with adversarial corruptions.


HybridRegretBoundsforCombinatorial Semi-BanditsandAdversarialLinearBandits

Neural Information Processing Systems

Theformer means that the algorithm will work nearly optimally in all environments in an adversarial setting, a stochastic setting, or a stochastic setting with adversarial corruptions.




Near-OptimalReinforcementLearningwithSelf-Play

Neural Information Processing Systems

This paper considers the problem of designing optimal algorithms for reinforcement learning in two-player zero-sum games. We focus on self-play algorithms which learn theoptimal policy by playing againstitself without any direct supervision.