Position: Benchmarking is Limited in Reinforcement Learning Research

Jordan, Scott M., White, Adam, da Silva, Bruno Castro, White, Martha, Thomas, Philip S.

arXiv.org Artificial Intelligence 

Many articles have pointed out problems Novel reinforcement learning algorithms, or improvements with reproducibility (Henderson et al., 2018; Islam et al., on existing ones, are commonly justified 2017; Smith, 2018; Engstrom et al., 2020) or statistical by evaluating their performance on benchmark analyses (Colas et al., 2018; Agarwal et al., 2021). Other environments and are compared to an everchanging works have examined methodological issues (Patterson et al., set of standard algorithms. However, 2023), showing that performance evaluation is sensitive to despite numerous calls for improvements, experimental subtle factors such as hyperparameter selection, score normalization, practices continue to produce misleading and the weight assigned to each task in an aggregate or unsupported claims. One reason for the ongoing performance measure (Jordan et al., 2020; Whiteson substandard practices is that conducting et al., 2011; Balduzzi et al., 2018; Eimer et al., 2023). These rigorous benchmarking experiments requires substantial works propose new methods to control for sources of variation computational time. This work investigates in performance and make the process more rigorous, the sources of increased computation costs including running more trials.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found