Schoenberg-Rao distances: Entropy-based and geometry-aware statistical Hilbert distances
Hadjeres, Gaëtan, Nielsen, Frank
Choosing a suitable statistical distance [11, 2] based on first principles is essential to ensure the relevancy and effectiveness of tasks in machine learning. Various statistical distances 1 have been proposed in the literature, starting from the early days of Mahalanobis [22] with his eponym distance. Later, these statistical distances have been studied under the umbrella of families of statistical distances called divergences: The Csiszár f-divergences [9] I f (p q) p(x)f(q(x)/p(x))dP (x) I f (q p) defined for a convex generator f(u) with f(1) 0 (including the total variation metric for f(u) u 1 and the Kullback-Leibler (KL) divergence for f(u) log u) with conjugate generator f (u) uf(1/u), the Bregman divergences [6], the Jensen divergences (also called Burbea-Rao divergences [7]), etc. From the viewpoint of statistical invariances, f-divergences are the only invariant 2 separable 3 divergences in information geometry [1]: The f-divergences are kept unchanged under a diffeomorphism of the sample space (and by reparameterization with the sufficient statistics) and under a smooth one-to-one mapping of the parameter space of parametric families of distributions [24].
Feb-19-2020