Goto

Collaborating Authors

 Reinforcement Learning


Constrainedepisodicreinforcementlearningin concave-convexandknapsacksettings

Neural Information Processing Systems

Our approach relies on the principle ofoptimism under uncertaintyto efficiently explore. Our learning algorithms optimizetheiractions withrespect toamodel based ontheempirical statistics, while optimistically overestimating rewards and underestimating the resource consumption (i.e., overestimating the distance from the constraint).






ALawofIteratedLogarithmforMulti-Agent ReinforcementLearning

Neural Information Processing Systems

In contrast, the mathematics needed to analyze such schemes is what forms the focus in Stochastic Approximation (SA) theory [2, 4]. More generally, SA refers to an iterative scheme that helps find zeroes or optimal points of a function, for which only noisy evaluationsarepossible.