Bayesian Inference
Quantifying the multi-objective cost of uncertainty
Yoon, Byung-Jun, Qian, Xiaoning, Dougherty, Edward R.
Investigating real-world systems and phenomena typically requires complex models that involve a large number of parameters. Even with sizeable amount of observation data, the high complexity of the model may render accurate parameter estimation impossible. While finding a reliable point estimate of the parameter vector may not be possible in such a case, it may be possible to identify the parameter ranges based on the available data and/or prior system knowledge, or in a more general setting, we may assume a joint distribution of the model parameters. Since different parameter values are possible, this gives rise to an uncertainty class of all possible models [1, 2]. Furthermore, this naturally places the model in a Bayesian framework, where the likelihood of every possible model in the uncertainty class is described by a prior that could be constructed from prior system knowledge and existing data [3, 4]. Given an uncertain model and its uncertainty class, how can one mathematically quantify the amount of uncertainty present in the model? Common approaches include estimating the variance or entropy of the uncertain parameters, as they both provide a simple and intuitive measure of the model uncertainty. However, they both have a critical downside from a practical perspective. In practical applications that involve mathematical modeling of a complex system, one cares about the model as it can serve as a vehicle for designing an effective operator (i.e., controller, classifier, filter) that can act on the system of interest or the data produced therefrom.
Ensembling geophysical models with Bayesian Neural Networks
Sengupta, Ushnish, Amos, Matt, Hosking, J. Scott, Rasmussen, Carl Edward, Juniper, Matthew, Young, Paul J.
Ensembles of geophysical models improve prediction accuracy and express uncertainties. We develop a novel data-driven ensembling strategy for combining geophysical models using Bayesian Neural Networks, which infers spatiotemporally varying model weights and bias, while accounting for heteroscedastic uncertainties in the observations. This produces more accurate and uncertaintyaware predictions without sacrificing interpretability. Applied to the prediction of total column ozone from an ensemble of 15 chemistry-climate models, we find that the Bayesian neural network ensemble (BayNNE) outperforms existing methods for ensembling physical models, achieving a 49.4% reduction in RMSE for temporal extrapolation, and a 67.4% reduction in RMSE for polar data voids, compared to a weighted mean. Uncertainty is also well-characterized, with 91.9% of the data points in our extrapolation validation dataset lying within 2 standard deviations and 98.9% within 3 standard deviations.
Exact Symbolic Inference in Probabilistic Programs via Sum-Product Representations
Saad, Feras A., Rinard, Martin C., Mansinghka, Vikash K.
We present the Sum-Product Probabilistic Language (SPPL), a new system that automatically delivers exact solutions to a broad range of probabilistic inference queries. SPPL symbolically represents the full distribution on execution traces specified by a probabilistic program using a generalization of sum-product networks. SPPL handles continuous and discrete distributions, many-to-one numerical transformations, and a query language that includes general predicates on random variables. We formalize SPPL in terms of a novel translation strategy from probabilistic programs to a semantic domain of sum-product representations, present new algorithms for exactly conditioning on and computing probabilities of queries, and prove their soundness under the semantics. We present techniques for improving the scalability of translation and inference by automatically exploiting conditional independences and repeated structure in SPPL programs. We implement a prototype of SPPL with a modular architecture and evaluate it on a suite of common benchmarks, which establish that our system is up to 3500x faster than state-of-the-art systems for fairness verification; up to 1000x faster than state-of-the-art symbolic algebra techniques; and can compute exact probabilities of rare events in milliseconds.
Data Driven Density Functional Theory: A case for Physics Informed Learning
Yatsyshin, Peter, Kalliadasis, Serafim, Duncan, Andrew B.
We propose a novel data-driven approach to solving a classical statistical mechanics problem: given data on collective motion of particles, characterise the set of free energies associated with the system of particles. We demonstrate empirically that the particle data contains all the information necessary to infer a free energy. While traditional physical modelling seeks to construct analytically tractable approximations, the proposed approach leverages modern Bayesian computational capabilities to accomplish this in a purely data-driven fashion. The Bayesian paradigm permits us to combine underpinning physical principles with simulation data to obtain uncertainty-quantified predictions of the free energy, in the form of a probability distribution over the family of free energies consistent with the observed particle data. In the present work we focus on classical statistical mechanical systems with excluded volume interactions. Using standard coarse-graining methods, our results can be made applicable to systems with realistic attractive-repulsive interactions. We validate our method on a paradigmatic and computationally cheap case of a one-dimensional fluid. With the appropriate particle data, it is possible to learn canonical and grand-canonical representations of the underlying physical system. Extensions to higher-dimensional systems are conceptually straightforward.
Effects of Model Misspecification on Bayesian Bandits: Case Studies in UX Optimization
Sweeney, Mack, van Adelsberg, Matthew, Laskey, Kathryn, Domeniconi, Carlotta
Bayesian bandits using Thompson Sampling have seen increasing success in recent years. Yet existing value models (of rewards) are misspecified on many real-world problem. We demonstrate this on the User Experience Optimization (UXO) problem, providing a novel formulation as a restless, sleeping bandit with unobserved confounders plus optional stopping. Our case studies show how common misspecifications can lead to sub-optimal rewards, and we provide model extensions to address these, along with a scientific model building process practitioners can adopt or adapt to solve their own unique problems. To our knowledge, this is the first study showing the effects of overdispersion on bandit explore/exploit efficacy, tying the common notions of under- and over-confidence to over- and under-exploration, respectively. We also present the first model to exploit cointegration in a restless bandit, demonstrating that finite regret and fast and consistent optional stopping are possible by moving beyond simpler windowing, discounting, and drift models.
Bayesian Distance Weighted Discrimination
Distance weighted discrimination (DWD) is a linear discrimination method that is particularly well-suited for classification tasks with high-dimensional data. The DWD coefficients minimize an intuitive objective function, which can solved very efficiently using state-of-the-art optimization techniques. However, DWD has not yet been cast into a model-based framework for statistical inference. In this article we show that DWD identifies the mode of a proper Bayesian posterior distribution, that results from a particular link function for the class probabilities and a shrinkage-inducing proper prior distribution on the coefficients. We describe a relatively efficient Markov chain Monte Carlo (MCMC) algorithm to simulate from the true posterior under this Bayesian framework. We show that the posterior is asymptotically normal and derive the mean and covariance matrix of its limiting distribution. Through several simulation studies and an application to breast cancer genomics we demonstrate how the Bayesian approach to DWD can be used to (1) compute well-calibrated posterior class probabilities, (2) assess uncertainty in the DWD coefficients and resulting sample scores, (3) improve power via semi-supervised analysis when not all class labels are available, and (4) automatically determine a penalty tuning parameter within the model-based framework. R code to perform Bayesian DWD is available at https://github.com/lockEF/BayesianDWD .
Fixing Asymptotic Uncertainty of Bayesian Neural Networks with Infinite ReLU Features
Kristiadi, Agustinus, Hein, Matthias, Hennig, Philipp
Approximate Bayesian methods can mitigate overconfidence in ReLU networks. However, far away from the training data, even Bayesian neural networks (BNNs) can still underestimate uncertainty and thus be overconfident. We suggest to fix this by considering an infinite number of ReLU features over the input domain that are never part of the training process and thus remain at prior values. Perhaps surprisingly, we show that this model leads to a tractable Gaussian process (GP) term that can be added to a pre-trained BNN's posterior at test time with negligible cost overhead. The BNN then yields structured uncertainty in the proximity of training data, while the GP prior calibrates uncertainty far away from them. As a key contribution, we prove that the added uncertainty yields cubic predictive variance growth, and thus the ideal uniform (maximum entropy) confidence in multi-class classification far from the training data. Calibrated uncertainty is crucial for safety-critical decision making by neural networks (NNs) (Amodei et al., 2016). Standard training methods of NNs yield point estimates that, even if they are highly accurate, can still be severely overconfident (Guo et al., 2017).
Recyclable Gaussian Processes
Moreno-Muรฑoz, Pablo, Artรฉs-Rodrรญguez, Antonio, รlvarez, Mauricio A.
We present a new framework for recycling independent variational approximations to Gaussian processes. The main contribution is the construction of variational ensembles given a dictionary of fitted Gaussian processes without revisiting any subset of observations. Our framework allows for regression, classification and heterogeneous tasks, i.e. mix of continuous and discrete variables over the same input domain. We exploit infinite-dimensional integral operators based on the Kullback-Leibler divergence between stochastic processes to re-combine arbitrary amounts of variational sparse approximations with different complexity, likelihood model and location of the pseudo-inputs. Extensive results illustrate the usability of our framework in large-scale distributed experiments, also compared with the exact inference models in the literature.
Self-Supervised Variational Auto-Encoders
Gatopoulos, Ioannis, Tomczak, Jakub M.
Density estimation, compression and data generation are crucial tasks in artificial intelligence. Variational Auto-Encoders (VAEs) constitute a single framework to achieve these goals. Here, we present a novel class of generative models, called self-supervised Variational Auto-Encoder (selfVAE), that utilizes deterministic and discrete variational posteriors. This class of models allows to perform both conditional and unconditional sampling, while simplifying the objective function. First, we use a single self-supervised transformation as a latent variable, where a transformation is either downscaling or edge detection. Next, we consider a hierarchical architecture, i.e., multiple transformations, and we show its benefits compared to the VAE. The flexibility of selfVAE in data reconstruction finds a particularly interesting use case in data compression tasks, where we can trade-off memory for better data quality, and vice-versa. We present performance of our approach on three benchmark image data (Cifar10, Imagenette64, and CelebA).
Sequential Changepoint Detection in Neural Networks with Checkpoints
Titsias, Michalis K., Sygnowski, Jakub, Chen, Yutian
We introduce a framework for online changepoint detection and simultaneous model learning which is applicable to highly parametrized models, such as deep neural networks. It is based on detecting changepoints across time by sequentially performing generalized likelihood ratio tests that require only evaluations of simple prediction score functions. This procedure makes use of checkpoints, consisting of early versions of the actual model parameters, that allow to detect distributional changes by performing predictions on future data. We define an algorithm that bounds the Type I error in the sequential testing procedure. We demonstrate the efficiency of our method in challenging continual learning applications with unknown task changepoints, and show improved performance compared to online Bayesian changepoint detection.