Exploration Bonus for Regret Minimization in Discrete and Continuous Average Reward MDPs
Jian QIAN, Ronan Fruit, Matteo Pirotta, Alessandro Lazaric
–Neural Information Processing Systems
The exploration bonus is an effective approach to manage the exploration-exploitation trade-off in Markov Decision Processes (MDPs).
Neural Information Processing Systems
Aug-20-2025, 05:45:10 GMT