Unsupervised Behavior Extraction via Random Intent Priors
–Neural Information Processing Systems
Reward-free data is abundant and contains rich prior knowledge of human behaviors, but it is not well exploited by offline reinforcement learning (RL) algorithms. In this paper, we propose UBER, an unsupervised approach to extract useful behaviors from offline reward-free datasets via diversified rewards.
Neural Information Processing Systems
Oct-9-2025, 03:14:47 GMT