BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices

May-26-2025, 19:12:22 GMT–Neural Information Processing Systems

AI models are increasingly prevalent in high-stakes environments, necessitating thorough assessment of their capabilities and risks. Benchmarks are popular for measuring these attributes and for comparing model performance, tracking progress, and identifying weaknesses in foundation and non-foundation models. They can inform model selection for downstream tasks and influence policy initiatives. However, not all benchmarks are the same: their quality depends on their design and usability. In this paper, we develop an assessment framework considering 40 best practices across a benchmark's life cycle and evaluate 25 AI benchmarks against it.

artificial intelligence, benchmark, machine learning, (6 more...)

Neural Information Processing Systems

May-26-2025, 19:12:22 GMT

Conferences Web Page

Add feedback

Technology:
- Information Technology > Artificial Intelligence > Machine Learning (0.44)