Goto

Collaborating Authors

 Country



Policy Improvement using Language Feedback Models

Neural Information Processing Systems

First, by using LFMs to identify desirable behaviour to imitate, we improve in task-completion rate over strong behavioural cloning baselines on three distinct language grounding environments (Touchdown, ScienceWorld, and ALFWorld). Second, imitation learning using LFMs outperform using LLMs as experts to directly predict actions, when controlling for the number of LLM output tokens.


HeterogeneousSkillLearningforMulti-agent Tasks

Neural Information Processing Systems

Meanwhile, diverseskill-based policies are generated through a novel skill-based policy learning method. To promote efficient skill discovery, a mutual information based intrinsic reward function is constructed.


Temporal Regularization for Markov Decision Process

Neural Information Processing Systems

Yetinreinforcementlearning,duetothenatureofthe Bellman equation, there isanopportunity toalsoexploit temporal regularization based on smoothness in value estimates over trajectories. This paper explores a class of methods for temporal regularization.





Exponentially Weighted Imitation Learning for Batched Historical Data

Neural Information Processing Systems

We consider deep policy learning with only batched historical trajectories. The main challenge of this problem is that the learner no longer has a simulator or "environment oracle" as in most reinforcement learning settings.


GT-GAN: GeneralPurposeTimeSeriesSynthesis withGenerativeAdversarialNetworks

Neural Information Processing Systems

However, there are no existing generative models that showgood performance for both types without anymodel changes. Therefore, we present a general purpose model capable of synthesizing regular and irregular time series data.