Goto

Collaborating Authors

 Technology



OfflineReinforcementLearningwithReverse Model-basedImagination

Neural Information Processing Systems

However, in many real-world applications, collecting sufficient exploratory interactions is usually impractical, because online datacollection canbecostlyorevendangerous, suchasinhealthcare [4]andautonomous driving [5]. To address this challenge, offline RL [6, 7] develops a new learning paradigm that trains RL agents only with pre-collected offline datasets and thus can abstract away from the cost of online exploration [8-17].


OfflineReinforcementLearningwithReverse Model-basedImagination

Neural Information Processing Systems

However, in many real-world applications, collecting sufficient exploratory interactions is usually impractical, because online datacollection canbecostlyorevendangerous, suchasinhealthcare [4]andautonomous driving [5]. To address this challenge, offline RL [6, 7] develops a new learning paradigm that trains RL agents only with pre-collected offline datasets and thus can abstract away from the cost of online exploration [8-17].


Symmetry-inducedDisentanglementonGraphs

Neural Information Processing Systems

Disentanglementhasbeen formalized using a symmetry-centric notion for unstructured spaces, however, graphs have eluded a similarly rigorous treatment. We fill this gap with a new notionofconditional symmetryfordisentanglement, andleveragetoolsfromLie algebras toencode graph properties intosubgroups using suitable adaptations of generative models such as Variational Autoencoders.




cc4d9cfc45325e460b455a820d5f212c-Supplemental-Conference.pdf

Neural Information Processing Systems

The observations are based onproprioception andonegocentricvision. The observations are based onproprioception, sword position and orientation, holeposition. Notetheseenvironmentsare related tothe domains that havebeen proposed for use inoffline RL benchmarks [Gulcehre etal., 2020]; however,the experiments we perform inthis work require availability ofthe expert policy, so we do not use offline data, but instead train new experts and perform experiments in the very low data regime. The main method we consider isAPC described inSection 3from main paper for offline experts cloning experiments. We can also consider resampling a new action, but we empirically found that cross-entropy worked better, see Appendix(8.6).




LearningwithNoisyCorrespondence forCross-modalMatching

Neural Information Processing Systems

In practice, however, such an assumption is extremely expensive even impossible to satisfy. Based on this observation, we reveal and study alatent and challenging direction in cross-modal matching, named noisy correspondence, which could be regarded as a new paradigm of noisylabels.