Goto

Collaborating Authors

 Country



ImprovingSampleComplexityBoundsfor(Natural) Actor-CriticAlgorithms

Neural Information Processing Systems

The goal of reinforcement learning (RL) [39] is to maximize the expected total reward by taking actions according toapolicyinastochastic environment, whichismodelled asaMarkovdecision process (MDP) [4]. To obtain an optimal policy, one popular method is the direct maximization of the expected total reward via gradient ascent, which is referred to as the policy gradient (PG) method [40,47].



Log-PolarSpaceConvolutionLayers: Appendix

Neural Information Processing Systems

The center pixelsofallareas form thecenter set. In this way, we obtain the correlation scores from all3 8 regions to the center pixel, as shown in Table A1. We run all models for only one time. Method Sum Max NoCenterConv Mean Acc. Effects ofweight regularization.InTab.A3,weevaluate theeffectsoftheweight regularization in Eq. (3) in the main text based on AlexNet.