Reviews: Adaptive Batch Size for Safe Policy Gradients
–Neural Information Processing Systems
Summary: This paper derives conditions for guaranteed improvement when using policy gradient methods. These conditions are for stochastic gradient estimates and also bound the amount of improvement with high probability. The authors then show how these bounds can be optimized by properly selecting the step size and batch size parameters. This is in contrast to previous work that only considers how the step size can be optimized. The result is an algorithm that can guarantee improvement with high probability.
Neural Information Processing Systems
Oct-8-2024, 12:29:27 GMT
- Technology: