Technology
OfflineReinforcementLearningwithReverse Model-basedImagination
However, in many real-world applications, collecting sufficient exploratory interactions is usually impractical, because online datacollection canbecostlyorevendangerous, suchasinhealthcare [4]andautonomous driving [5]. To address this challenge, offline RL [6, 7] develops a new learning paradigm that trains RL agents only with pre-collected offline datasets and thus can abstract away from the cost of online exploration [8-17].
OfflineReinforcementLearningwithReverse Model-basedImagination
However, in many real-world applications, collecting sufficient exploratory interactions is usually impractical, because online datacollection canbecostlyorevendangerous, suchasinhealthcare [4]andautonomous driving [5]. To address this challenge, offline RL [6, 7] develops a new learning paradigm that trains RL agents only with pre-collected offline datasets and thus can abstract away from the cost of online exploration [8-17].
Symmetry-inducedDisentanglementonGraphs
Disentanglementhasbeen formalized using a symmetry-centric notion for unstructured spaces, however, graphs have eluded a similarly rigorous treatment. We fill this gap with a new notionofconditional symmetryfordisentanglement, andleveragetoolsfromLie algebras toencode graph properties intosubgroups using suitable adaptations of generative models such as Variational Autoencoders.
cc4d9cfc45325e460b455a820d5f212c-Supplemental-Conference.pdf
The observations are based onproprioception andonegocentricvision. The observations are based onproprioception, sword position and orientation, holeposition. Notetheseenvironmentsare related tothe domains that havebeen proposed for use inoffline RL benchmarks [Gulcehre etal., 2020]; however,the experiments we perform inthis work require availability ofthe expert policy, so we do not use offline data, but instead train new experts and perform experiments in the very low data regime. The main method we consider isAPC described inSection 3from main paper for offline experts cloning experiments. We can also consider resampling a new action, but we empirically found that cross-entropy worked better, see Appendix(8.6).