Goto

Collaborating Authors

 Statistical Learning





Cross-validation Confidence Intervals for Test Error Pierre Bayle

Neural Information Processing Systems

This work develops central limit theorems for cross-validation and consistent estimators of its asymptotic variance under weak stability conditions on the learning algorithm. Together, these results provide practical, asymptotically-exact confidence intervals for k -fold test error and valid, powerful hypothesis tests of whether one learning algorithm has smaller k -fold test error than another. These results are also the first of their kind for the popular choice of leave-one-out cross-validation. In our real-data experiments with diverse learning algorithms, the resulting intervals and tests outperform the most popular alternative methods from the literature.


We thank all reviewers for their time and feedback; we address common and individual comments in turn

Neural Information Processing Systems

We thank all reviewers for their time and feedback; we address common and individual comments in turn. For example, we implemented the ridge regression CI from [Thm. K Figure 1, our maximum width is 1.03, but the Hold even when training error is a poor proxy for test error due to overfitting (e.g., 1-nearest neighbor has training In the revision, we will clarify the formal definitions of "asymptotically exact" (coverage converging In the revision, we will highlight that (3.2) in Thm. 2 already implies a "non-45 F & G for detailed examples of simple learning problems excluded by past work but covered by ours).