Learning Graphical Models
Neural Approximate Sufficient Statistics for Implicit Models
Chen, Yanzhi, Zhang, Dinghuai, Gutmann, Michael, Courville, Aaron, Zhu, Zhanxing
We consider the fundamental problem of how to automatically construct summary statistics for implicit generative models where the evaluation of likelihood function is intractable but sampling / simulating data from the model is possible. The idea is to frame the task of constructing sufficient statistics as learning mutual information maximizing representation of the data. This representation is computed by a deep neural network trained by a joint statistic-posterior learning strategy. We apply our approach to both traditional approximate Bayesian computation (ABC) and recent neural likelihood approaches, boosting their performance on a range of tasks.
How to Explain Key Machine Learning Algorithms at an Interview - KDnuggets
Linear Regression involves finding a'line of best fit' that represents a dataset using the least squares method. The least squares method involves finding a linear equation that minimizes the sum of squared residuals. A residual is equal to the actual minus predicted value. To give an example, the red line is a better line of best fit than the green line because it is closer to the points, and thus, the residuals are smaller. Ridge regression, also known as L2 Regularization, is a regression technique that introduces a small amount of bias to reduce overfitting.
Analytics play a critical role in digital transformation. Here's how to start.
Artificial intelligence (AI), machine learning (ML), and predictive analytics (PA) are current buzzwords eclipsing the tech industry hype curve. At this point, most tech savvy people are tired of hearing how these techniques will save the day. Let's eliminate some of the hype and paint a more realistic picture of what's happening with these technologies. All three are surprisingly old concepts. Predictive analytics can trace its origins back to Thomas Bayes, who laid the foundation to Bayesian probability theory in 1736. Nothing new here, except within the last 50 years, computers have made it much easier to work with the math.
Can I Trust My Fairness Metric? Assessing Fairness with Unlabeled Data and Bayesian Inference
Ji, Disi, Smyth, Padhraic, Steyvers, Mark
We investigate the problem of reliably assessing group fairness when labeled examples are few but unlabeled examples are plentiful. We propose a general Bayesian framework that can augment labeled data with unlabeled data to produce more accurate and lower-variance estimates compared to methods based on labeled data alone. Our approach estimates calibrated scores for unlabeled examples in each group using a hierarchical latent variable model conditioned on labeled examples. This in turn allows for inference of posterior distributions with associated notions of uncertainty for a variety of group fairness metrics. We demonstrate that our approach leads to significant and consistent reductions in estimation error across multiple well-known fairness datasets, sensitive attributes, and predictive models. The results show the benefits of using both unlabeled data and Bayesian inference in terms of assessing whether a prediction model is fair or not.
Multiple-view clustering for correlation matrices based on Wishart mixture model
Tokuda, Tomoki, Yamashita, Okito, Yoshimoto, Junichiro
A multiple-view clustering method is a powerful analytical tool for high-dimensional data, such as functional magnetic resonance imaging (fMRI). It can identify clustering patterns of subjects depending on their functional connectivity in specific brain areas. However, when one applies an existing method to fMRI data, there is a need to simplify the data structure, independently dealing with elements in a functional connectivity matrix, that is, a correlation matrix. In general, elements in a correlation matrix are closely associated. Hence, such a simplification may distort the clustering results. To overcome this problem, we propose a novel multiple-view clustering method based on the Wishart mixture model, which preserves the correlation matrix structure. The uniqueness of this method is that the multiple-view clustering of subjects is based on particular networks of nodes (or regions of interest (ROIs) in fMRI), optimized in a data-driven manner. Hence, it can identify multiple underlying pairs of associations between a subject cluster solution and a ROI network. The key assumption of the method is independence among networks, which is effectively addressed by whitening correlation matrices. We applied the proposed method to synthetic and fMRI data, demonstrating the usefulness and power of the proposed method.
A Contour Stochastic Gradient Langevin Dynamics Algorithm for Simulations of Multi-modal Distributions
Deng, Wei, Lin, Guang, Liang, Faming
We propose an adaptively weighted stochastic gradient Langevin dynamics algorithm (SGLD), so-called contour stochastic gradient Langevin dynamics (CSGLD), for Bayesian learning in big data statistics. The proposed algorithm is essentially a \emph{scalable dynamic importance sampler}, which automatically \emph{flattens} the target distribution such that the simulation for a multi-modal distribution can be greatly facilitated. Theoretically, we prove a stability condition and establish the asymptotic convergence of the self-adapting parameter to a {\it unique fixed-point}, regardless of the non-convexity of the original energy function; we also present an error analysis for the weighted averaging estimators. Empirically, the CSGLD algorithm is tested on multiple benchmark datasets including CIFAR10 and CIFAR100. The numerical results indicate its superiority over the existing state-of-the-art algorithms in training deep neural networks.
ABC-Di: Approximate Bayesian Computation for Discrete Data
Auzina, Ilze Amanda, Tomczak, Jakub M.
Many real-life problems are represented as a black-box, i.e., the internal workings are inaccessible or a closed-form mathematical expression of the likelihood function cannot be defined. For continuous random variables likelihood-free inference problems can be solved by a group of methods under the name of Approximate Bayesian Computation (ABC). However, a similar approach for discrete random variables is yet to be formulated. Here, we aim to fill this research gap. We propose to use a population-based MCMC ABC framework. Further, we present a valid Markov kernel, and propose a new kernel that is inspired by Differential Evolution. We assess the proposed approach on a problem with the known likelihood function, namely, discovering the underlying diseases based on a QMR-DT Network, and three likelihood-free inference problems: (i) the QMR-DT Network with the unknown likelihood function, (ii) learning binary neural network, and (iii) Neural Architecture Search. The obtained results indicate the high potential of the proposed framework and the superiority of the new Markov kernel.
Faster Convergence of Stochastic Gradient Langevin Dynamics for Non-Log-Concave Sampling
Zou, Difan, Xu, Pan, Gu, Quanquan
We establish a new convergence analysis of stochastic gradient Langevin dynamics (SGLD) for sampling from a class of distributions that can be non-log-concave. At the core of our approach is a novel conductance analysis of SGLD using an auxiliary time-reversible Markov Chain. Under certain conditions on the target distribution, we prove that $\tilde O(d^4\epsilon^{-2})$ stochastic gradient evaluations suffice to guarantee $\epsilon$-sampling error in terms of the total variation distance, where $d$ is the problem dimension, which improves existing results on the convergence rate of SGLD (Raginsky et al., 2017; Xu et al., 2018). We further show that provided an additional Hessian Lipschitz condition on the log-density function, SGLD is guaranteed to achieve $\epsilon$-sampling error within $\tilde O(d^{15/4}\epsilon^{-3/2})$ stochastic gradient evaluations. Our proof technique provides a new way to study the convergence of Langevin based algorithms, and sheds some light on the design of fast stochastic gradient based sampling algorithms.
Learning Exponential Family Graphical Models with Latent Variables using Regularized Conditional Likelihood
Taeb, Armeen, Shah, Parikshit, Chandrasekaran, Venkat
Fitting a graphical model to a collection of random variables given sample observations is a challenging task if the observed variables are influenced by latent variables, which can induce significant confounding statistical dependencies among the observed variables. We present a new convex relaxation framework based on regularized conditional likelihood for latent-variable graphical modeling in which the conditional distribution of the observed variables conditioned on the latent variables is given by an exponential family graphical model. In comparison to previously proposed tractable methods that proceed by characterizing the marginal distribution of the observed variables, our approach is applicable in a broader range of settings as it does not require knowledge about the specific form of distribution of the latent variables and it can be specialized to yield tractable approaches to problems in which the observed data are not well-modeled as Gaussian. We demonstrate the utility and flexibility of our framework via a series of numerical experiments on synthetic as well as real data.
Bayesian Inference for Optimal Transport with Stochastic Cost
Mallasto, Anton, Heinonen, Markus, Kaski, Samuel
In machine learning and computer vision, optimal transport has had significant success in learning generative models and defining metric distances between structured and stochastic data objects, that can be cast as probability measures. The key element of optimal transport is the so called lifting of an \emph{exact} cost (distance) function, defined on the sample space, to a cost (distance) between probability measures over the sample space. However, in many real life applications the cost is \emph{stochastic}: e.g., the unpredictable traffic flow affects the cost of transportation between a factory and an outlet. To take this stochasticity into account, we introduce a Bayesian framework for inferring the optimal transport plan distribution induced by the stochastic cost, allowing for a principled way to include prior information and to model the induced stochasticity on the transport plans. Additionally, we tailor an HMC method to sample from the resulting transport plan posterior distribution.