Statistical Learning
Posterior Sampling Based on Gradient Flows of the MMD with Negative Distance Kernel
Hagemann, Paul, Hertrich, Johannes, Altekrüger, Fabian, Beinert, Robert, Chemseddine, Jannis, Steidl, Gabriele
We propose conditional flows of the maximum mean discrepancy (MMD) with the negative distance kernel for posterior sampling and conditional generative modeling. This MMD, which is also known as energy distance, has several advantageous properties like efficient computation via slicing and sorting. We approximate the joint distribution of the ground truth and the observations using discrete Wasserstein gradient flows and establish an error bound for the posterior distributions. Further, we prove that our particle flow is indeed a Wasserstein gradient flow of an appropriate functional. The power of our method is demonstrated by numerical examples including conditional image generation and inverse problems like superresolution, inpainting and computed tomography in low-dose and limited-angle settings.
Hoeffding's Inequality for Markov Chains under Generalized Concentrability Condition
Chen, Hao, Gupta, Abhishek, Sun, Yin, Shroff, Ness
This paper studies Hoeffding's inequality for Markov chains under the generalized concentrability condition defined via integral probability metric (IPM). The generalized concentrability condition establishes a framework that interpolates and extends the existing hypotheses of Markov chain Hoeffding-type inequalities. The flexibility of our framework allows Hoeffding's inequality to be applied beyond the ergodic Markov chains in the traditional sense. We demonstrate the utility by applying our framework to several non-asymptotic analyses arising from the field of machine learning, including (i) a generalization bound for empirical risk minimization with Markovian samples, (ii) a finite sample guarantee for Ployak-Ruppert averaging of SGD, and (iii) a new regret bound for rested Markovian bandits with general state space. Keywords: Hoeffding's inequality, Markov chains, Dobrushin coefficient, integral probability metric, concentration of measures, ergodicity, empirical risk minimization, stochastic gradient descent, rested bandit
Beyond Stationarity: Convergence Analysis of Stochastic Softmax Policy Gradient Methods
Klein, Sara, Weissmann, Simon, Döring, Leif
Markov Decision Processes (MDPs) are a formal framework for modeling and solving sequential decision-making problems. In finite-time horizons such problems are relevant for instance for optimal stopping or specific supply chain problems, but also in the training of large language models. In contrast to infinite horizon MDPs optimal policies are not stationary, policies must be learned for every single epoch. In practice all parameters are often trained simultaneously, ignoring the inherent structure suggested by dynamic programming. This paper introduces a combination of dynamic programming and policy gradient called dynamic policy gradient, where the parameters are trained backwards in time. For the tabular softmax parametrisation we carry out the convergence analysis for simultaneous and dynamic policy gradient towards global optima, both in the exact and sampled gradient settings without regularisation. It turns out that the use of dynamic policy gradient training much better exploits the structure of finite-time problems which is reflected in improved convergence bounds.
High-dimensional SGD aligns with emerging outlier eigenspaces
Arous, Gerard Ben, Gheissari, Reza, Huang, Jiaoyang, Jagannath, Aukosh
We rigorously study the joint evolution of training dynamics via stochastic gradient descent (SGD) and the spectra of empirical Hessian and gradient matrices. We prove that in two canonical classification tasks for multi-class high-dimensional mixtures and either 1 or 2-layer neural networks, the SGD trajectory rapidly aligns with emerging low-rank outlier eigenspaces of the Hessian and gradient matrices. Moreover, in multi-layer settings this alignment occurs per layer, with the final layer's outlier eigenspace evolving over the course of training, and exhibiting rank deficiency when the SGD converges to sub-optimal classifiers. This establishes some of the rich predictions that have arisen from extensive numerical studies in the last decade about the spectra of Hessian and information matrices over the course of training in overparametrized networks.
Scaling Laws for Associative Memories
Cabannes, Vivien, Dohmatob, Elvis, Bietti, Alberto
Learning arguably involves the discovery and memorization of abstract rules. The aim of this paper is to study associative memory mechanisms. Our model is based on high-dimensional matrices consisting of outer products of embeddings, which relates to the inner layers of transformer language models. We derive precise scaling laws with respect to sample size and parameter size, and discuss the statistical efficiency of different estimators, including optimization-based algorithms. We provide extensive numerical experiments to validate and interpret theoretical results, including fine-grained visualizations of the stored memory associations.
Stationarity without mean reversion: Improper Gaussian process regression and improper kernels
Gaussian processes (GP) regression has gained substantial popularity in machine learning applications. The behavior of a GP regression depends on the choice of covariance function. Stationary covariance functions are favorite in machine learning applications. However, (non-periodic) stationary covariance functions are always mean reverting and can therefore exhibit pathological behavior when applied to data that does not relax to a fixed global mean value. In this paper, we show that it is possible to use improper GP prior with infinite variance to define processes that are stationary but not mean reverting. To this aim, we introduce a large class of improper kernels that can only be defined in this improper regime. Specifically, we introduce the Smooth Walk kernel, which produces infinitely smooth samples, and a family of improper Mat\'ern kernels, which can be defined to be $j$-times differentiable for any integer $j$. The resulting posterior distributions can be computed analytically and it involves a simple correction of the usual formulas. By analyzing both synthetic and real data, we demonstrate that these improper kernels solve some known pathologies of mean reverting GP regression while retaining most of the favourable properties of ordinary smooth stationary kernels.
Probabilistic Block Term Decomposition for the Modelling of Higher-Order Arrays
Hinrich, Jesper Løve, Mørup, Morten
Tensors or multi-way arrays naturally occur in practically all areas of science including psychology (i.e., human responses to questionnaire data according to scoring criteria of different objects), chemometrics (i.e., excitation and emission spectra across samples), biology (i.e., genetic expression of cell proles across time and experimental conditions), and knowledge representations (i.e., entity-entity relationships across predicates), see also [1] and references therein. To analyze these multi-way arrays accounting for their higher order structure tensor decompositions have become important tools to characterize and discover structure in these data, see [2, 1] for details. Tensor decompositions have historically focused on maximum likelihood estimation methods to obtain a point estimate to decompose the data, most predominately based on Gaussian likelihood (least squares estimation). Recently, there has been a rise in the development of Bayesian inference for tensor data, initially focusing on binary or count data, but now applied more broadly to various types of data, for an overview see [3, 4]. The benets of a Bayesian approach are that it characterizes the decomposition solution as a distribution, the so-called posterior distribution, which allows characterization of the uncertainty whereas priors acts as regularizers adding robustness and preventing issues of degeneracy. Additionally, it provides a principled way to incorporate a priori information. For a review on maximum likelihood based and Bayesian tensor decomposition, see [2] and [3], respectively. The two most common tensor decomposition methods are the Canonical Polyadic Decomposition/PARAFAC (CPD) and Tucker model. The CPD model represents the data through a sum of outer product rank-1 terms (i.e., separate multi-linear structures), whereas Tucker uses a multi-linear rank decomposition (i.e., with "connected" multi-linear structures).
Neural Bayes Estimators for Irregular Spatial Data using Graph Neural Networks
Sainsbury-Dale, Matthew, Richards, Jordan, Zammit-Mangion, Andrew, Huser, Raphaël
Neural Bayes estimators are neural networks that approximate Bayes estimators in a fast and likelihood-free manner. They are appealing to use with spatial models and data, where estimation is often a computational bottleneck. However, neural Bayes estimators in spatial applications have, to date, been restricted to data collected over a regular grid. These estimators are also currently dependent on a prescribed set of spatial locations, which means that the neural network needs to be re-trained for new data sets; this renders them impractical in many applications and impedes their widespread adoption. In this work, we employ graph neural networks to tackle the important problem of parameter estimation from data collected over arbitrary spatial locations. In addition to extending neural Bayes estimation to irregular spatial data, our architecture leads to substantial computational benefits, since the estimator can be used with any arrangement or number of locations and independent replicates, thus amortising the cost of training for a given spatial model. We also facilitate fast uncertainty quantification by training an accompanying neural Bayes estimator that approximates a set of marginal posterior quantiles. We illustrate our methodology on Gaussian and max-stable processes. Finally, we showcase our methodology in a global sea-surface temperature application, where we estimate the parameters of a Gaussian process model in 2,161 regions, each containing thousands of irregularly-spaced data points, in just a few minutes with a single graphics processing unit.
Generative Sliced MMD Flows with Riesz Kernels
Hertrich, Johannes, Wald, Christian, Altekrüger, Fabian, Hagemann, Paul
Maximum mean discrepancy (MMD) flows suffer from high computational costs in large scale computations. In this paper, we show that MMD flows with Riesz kernels $K(x,y) = - \Vert x-y\Vert^r$, $r \in (0,2)$ have exceptional properties which allow their efficient computation. We prove that the MMD of Riesz kernels, which is also known as energy distance, coincides with the MMD of their sliced version. As a consequence, the computation of gradients of MMDs can be performed in the one-dimensional setting. Here, for $r=1$, a simple sorting algorithm can be applied to reduce the complexity from $O(MN+N^2)$ to $O((M+N)\log(M+N))$ for two measures with $M$ and $N$ support points. As another interesting follow-up result, the MMD of compactly supported measures can be estimated from above and below by the Wasserstein-1 distance. For the implementations we approximate the gradient of the sliced MMD by using only a finite number $P$ of slices. We show that the resulting error has complexity $O(\sqrt{d/P})$, where $d$ is the data dimension. These results enable us to train generative models by approximating MMD gradient flows by neural networks even for image applications. We demonstrate the efficiency of our model by image generation on MNIST, FashionMNIST and CIFAR10.