Exploration Bonus for Regret Minimization in Discrete and Continuous Average Reward MDPs
Jian QIAN, Ronan Fruit, Matteo Pirotta, Alessandro Lazaric
–Neural Information Processing Systems
The exploration bonus is an effective approach to manage the explorationexploitation trade-offinMarkovDecision Processes (MDPs).
Neural Information Processing Systems
Feb-14-2026, 15:23:11 GMT
- Country:
- Asia > Middle East
- Jordan (0.04)
- North America
- Canada > British Columbia
- United States (0.04)
- Asia > Middle East
- Technology: