Kernel Two-Sample Tests in High Dimension: Interplay Between Moment Discrepancy and Dimension-and-Sample Orders

Yan, Jian, Zhang, Xianyang

arXiv.org Machine Learning 

Nonparametric two-sample testing, aiming to determine whether two collections of samples are from the same distribution without specifying the exact parametric forms of the distributions, is one of the fundamental problems in statistics. Such tests have found applications in various areas such as bioinformatics, anomaly detection, model criticism, audio and image processing. The most traditional tools in this domain include the Kolmogorov-Smirnov test [Smirnov (1939)], Cramér-von Mises test [Anderson (1962)], Wald-Wolfowitz runs test [Wald and Wolfowitz (1940)], and Wilcoxon-Mann-Whitney test [Mann and Whitney (1947)]. Extensions and generalizations of these tests were studied in Darling (1957), Weiss (1960), Bickel (1969), Friedman and Rafsky (1979), among many others. Modern nonparametric tests have been developed based on integral probability metrics [Sriperumbudur et al. (2012)]. Notable members include the kernel Maximum Mean Discrepancy (MMD) two-sample test [Gretton et al. (2012)] and energy distance-based two-sample test [Székely et al. (2004)]. These metrics are gaining increasing popularity in both the statistics and machine learning communities and they have been applied to a plethora of statistical problems including goodness-of-fit testing [Székely and Rizzo (2005)], nonparametric analysis of variance [Rizzo and Székely (2010)], change-point detection [Matteson and James (2014); Chakraborty and Zhang (2021)], finding representative points of a distribution [Mak and Joseph (2018)] and controlled variable selection [Romano et al. (2020)].