ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation

Liu, Qin, Dineen, Jacob, Huang, Yuxi, Zhang, Sheng, Poon, Hoifung, Zhou, Ben, Chen, Muhao

arXiv.org Artificial Intelligence 

Benchmarks are central to measuring the capabilities of large language models and guiding model development, yet widespread data leakage from pretraining corpora undermines their validity. Models can match memorized content rather than demonstrate true generalization, which inflates scores, distorts cross-model comparisons, and misrepresents progress. The process runs iteratively with in-context demonstrations that steer generation toward more challenging and diagnostic cases. The framework provides a scalable path to continuously evolve benchmarks in step with the rapid progress of foundation models. Benchmarks are indispensable for assessing large language models (LLM) capabilities and steering model development (Cobbe et al., 2021; Sakaguchi et al., 2021; Talmor et al., 2018; Hendrycks et al., 2020; Srivastava et al., 2023; Liang et al., 2022). Y et growing evidence that widely used benchmarks are partially or fully present in pretraining corpora of models poses a fundamental validity threat: models can exploit memorized content rather than demonstrating true generalization (Wu et al., 2025; Liang et al., 2025; Xu et al., 2024b; Dong et al., 2024; Balloccu et al., 2024; Jiang et al., 2024b). This pervasive data leakage fundamentally undermines the reliability of evaluation, causing inflated reporting scores, distorted cross-model comparisons, and misrepresented progress of development.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found