Goto

Collaborating Authors

 Technology







Checklist

Neural Information Processing Systems

When drawing 1,000 z from the priorp(z)of the latent space learned by PLAS, only4%of the samples are decoded as high-return actions, while inLAPO,45%ofthedecoded actions arehigh-return actions.


LAPO: Latent-VariableAdvantage-WeightedPolicy OptimizationforOfflineReinforcementLearning

Neural Information Processing Systems

But in practice, it requires querying the behavior policy which is unknown, and using an erroneous approximation of the behavior policy can negatively affect the performance ([39]).