Typical examples are hyperparameters selection [5, 38, 17, 6], data augmentation [11, 42], implicit deep learning [3] or neural architecture search [33]. Figure 1: Convergence curves of the two proposed methods on a toy problem.
However, it is often argued that correct predictions in the tail are more "interesting" or "rewarding," but the community has not yet settled on a metric capturing this intuitive concept.