Spectral Entry-wise Matrix Estimation for Low-Rank Reinforcement Learning

Jan-20-2025, 02:36:30 GMT–Neural Information Processing Systems

We study matrix estimation problems arising in reinforcement learning with low-rank structure. In low-rank bandits, the matrix to be recovered specifies the expected arm rewards, and for low-rank Markov Decision Processes (MDPs), it characterizes the transition kernel of the MDP. In both cases, each entry of the matrix carries important information, and we seek estimation methods with low entry-wise prediction error. Importantly, these methods further need to accommodate for inherent correlations in the available data (e.g. for MDPs, the data consists of system trajectories). We investigate the performance of simple spectral-based matrix estimation approaches: we show that they efficiently recover the singular subspaces of the matrix and exhibit nearly-minimal entry-wise prediction error.

algorithm, low-rank reinforcement learning, spectral entry-wise matrix estimation, (3 more...)

Neural Information Processing Systems

Jan-20-2025, 02:36:30 GMT

Conferences Web Page

Add feedback

Technology:
- Information Technology > Artificial Intelligence > Machine Learning > Reinforcement Learning (0.84)