Achieving Constant Regret in Linear Markov Decision Processes

Oct-10-2025, 20:29:34 GMT–Neural Information Processing Systems

We study the constant regret guarantees in reinforcement learning (RL). Our objective is to design an algorithm that incurs only finite regret over infinite episodes with high probability. We introduce an algorithm, Cert-LSVI-UCB, for misspec-ified linear Markov decision processes (MDPs) where both the transition kernel and the reward function can be approximated by some linear function up to mis-specification level ζ . At the core of Cert-LSVI-UCB is an innovative certified estimator, which facilitates a fine-grained concentration analysis for multi-phase value-targeted regression, enabling us to establish an instance-dependent regret bound that is constant w.r.t. the number of episodes.

algorithm, inequality hold, international conference, (13 more...)

Neural Information Processing Systems

Oct-10-2025, 20:29:34 GMT

Conferences PDF

Add feedback

Country:
- North America > United States
  - North Carolina > Orange County
    - Chapel Hill (0.04)
  - Massachusetts > Middlesex County
    - Cambridge (0.14)
  - California > Los Angeles County
    - Los Angeles (0.28)
- Europe > United Kingdom
  - England > Cambridgeshire > Cambridge (0.04)

Genre:
- Research Report > Experimental Study (0.92)

Industry:
- Health & Medicine (0.54)

Technology:
- Information Technology > Artificial Intelligence > Machine Learning
  - Reinforcement Learning (1.00)
  - Learning Graphical Models > Undirected Networks
    - Markov Models (0.70)

Duplicate Docs Excel Report

Title
Achieving Constant Regret in Linear Markov Decision Processes

Similar Docs Excel Report more

Title	Similarity	Source
None found