Regret Bounds for Learning State Representations in Reinforcement Learning

Ronald Ortner, Matteo Pirotta, Alessandro Lazaric, Ronan Fruit, Odalric-Ambrym Maillard

Jan-25-2025, 23:58:54 GMT–Neural Information Processing Systems

We consider the problem of online reinforcement learning when several state representations (mapping histories to a discrete state space) are available to the learning agent. At least one of these representations is assumed to induce a Markov decision process (MDP), and the performance of the agent is measured in terms of cumulative regret against the optimal policy giving the highest average reward in this MDP representation.

machine learning, markov model, reinforcement learning, (18 more...)

Neural Information Processing Systems

Jan-25-2025, 23:58:54 GMT

Conferences PDF

Add feedback

Country:
- Europe (0.46)
- North America > United States (0.28)

Technology:
- Information Technology > Artificial Intelligence > Machine Learning
  - Learning Graphical Models > Undirected Networks
    - Markov Models (0.39)
  - Reinforcement Learning (0.86)