A Meta-Analysis of Overfitting in Machine Learning

Dec-26-2025, 03:11:44 GMT–Neural Information Processing Systems

We conduct the first large meta-analysis of overfitting due to test set reuse in the machine learning community. Our analysis is based on over one hundred machine learning competitions hosted on the Kaggle platform over the course of several years. In each competition, numerous practitioners repeatedly evaluated their progress against a holdout set that forms the basis of a public ranking available throughout the competition. Performance on a separate test set used only once determined the final ranking. By systematically comparing the public ranking with the final ranking, we assess how much participants adapted to the holdout set over the course of a competition.

competition, meta-analysis, overfitting, (6 more...)

Neural Information Processing Systems

Dec-26-2025, 03:11:44 GMT

Conferences Web Page

Add feedback

Technology:
- Information Technology > Artificial Intelligence > Machine Learning (1.00)