Goto

Collaborating Authors

 Reinforcement Learning


DiffTORI: Differentiable Trajectory Optimization for Deep Reinforcement and Imitation Learning Weikang Wan

Neural Information Processing Systems

This paper introduces DiffTORI, which utilizes Diff erentiable T rajectory O ptimization as the policy representation to generate actions for deep R einforcement and I mitation learning. Trajectory optimization is a powerful and widely used algorithm in control, parameterized by a cost and a dynamics function.


CQM: Curriculum Reinforcement Learning with a Quantized World Model

Neural Information Processing Systems

Recent curriculum Reinforcement Learning (RL) has shown notable progress in solving complex tasks by proposing sequences of surrogate tasks. However, the previous approaches often face challenges when they generate curriculum goals in a high-dimensional space.








OptimisticCriticReconstructionandConstrained Fine-TuningforGeneralOffline-to-OnlineRL

Neural Information Processing Systems

Afterobtaining an optimistic and and aligned critic, we perform constrained fine-tuning to combat distribution shift during online learning.