Goto

Collaborating Authors

 Reinforcement Learning




Dynamic Regret of Adversarial Linear Mixture MDPs

Neural Information Processing Systems

We study reinforcement learning in episodic inhomogeneous MDPs with adversarial full-information rewards and the unknown transition kernel. We consider the linear mixture MDPs whose transition kernel is a linear mixture model and choose the dynamic regret as the performance measure.



The Value of Reward Lookahead in Reinforcement Learning

Neural Information Processing Systems

In reinforcement learning (RL), agents sequentially interact with changing environments while aiming to maximize the obtained rewards.