Position: Benchmarking is Limited in Reinforcement Learning Research
Jordan, Scott M., White, Adam, da Silva, Bruno Castro, White, Martha, Thomas, Philip S.
–arXiv.org Artificial Intelligence
Many articles have pointed out problems Novel reinforcement learning algorithms, or improvements with reproducibility (Henderson et al., 2018; Islam et al., on existing ones, are commonly justified 2017; Smith, 2018; Engstrom et al., 2020) or statistical by evaluating their performance on benchmark analyses (Colas et al., 2018; Agarwal et al., 2021). Other environments and are compared to an everchanging works have examined methodological issues (Patterson et al., set of standard algorithms. However, 2023), showing that performance evaluation is sensitive to despite numerous calls for improvements, experimental subtle factors such as hyperparameter selection, score normalization, practices continue to produce misleading and the weight assigned to each task in an aggregate or unsupported claims. One reason for the ongoing performance measure (Jordan et al., 2020; Whiteson substandard practices is that conducting et al., 2011; Balduzzi et al., 2018; Eimer et al., 2023). These rigorous benchmarking experiments requires substantial works propose new methods to control for sources of variation computational time. This work investigates in performance and make the process more rigorous, the sources of increased computation costs including running more trials.
arXiv.org Artificial Intelligence
Jun-23-2024
- Country:
- North America
- Canada > Alberta (0.14)
- United States
- New York (0.04)
- Massachusetts > Hampshire County
- Amherst (0.04)
- Europe
- Asia > Middle East
- Jordan (0.25)
- North America
- Genre:
- Research Report > New Finding (1.00)
- Industry:
- Leisure & Entertainment (0.67)
- Technology: