Stochastic Gradient Hamiltonian Monte Carlo with Variance Reduction for Bayesian Inference

Li, Zhize, Zhang, Tianyi, Li, Jian

arXiv.org Machine Learning 

Gradient-based Monte Carlo algorithms are useful tools for sampling posterior distributions. Similar to gradient descent algorithms, gradient-based Monte Carlo generates posterior samples iteratively using the gradient of loglikelihood. Langevin dynamics (LD) and Hamiltonian Monte Carlo (HMC) [Duane et al., 1987, Neal et al., 2011] are two important examples of gradient-based Monte Carlo sampling algorithms that are widely used in Bayesian inference. Since calculating likelihood on large datasets is expensive, people use stochastic gradients [Robbins and Monro, 1951] in place of full gradient, and have, for both Langevin dynamics and Hamiltonian Monte Carlo, developed their stochastic gradient counterparts[Welling and Teh, 2011, Chen et al., 2014]. Stochastic gradient Hamiltonian Monte Carlo (SGHMC) usually converges faster than stochastic gradient Langevin dynamics (SGLD) in practical machine learning tasks like covariance estimation of bivariate Gaussian and Bayesian neural networks for classification on MNIST dataset, as demonstrated in [Welling and Teh, 2011]. Similar phenomenon was also observed in [Chen et al., 2015] where SGHMC and SGLD were compared on both synthetic and real-world datasets. Intuitively speaking, comparing against SGLD, SGHMC has a momentum term that enables it to explore the parameter space of posterior distribution much faster when the gradient of log-likelihood becomes smaller. Very recently, [Dubey et al., 2016] borrowed the standard variance reduction techniques from the stochastic optimization literature [Johnson and Zhang, 2013, Defazio et al., 2014] and applied them on SGLD to obtain two variancereduced SGLD algorithms (called SAGA-LD and SVRG-LD) with improved theoretical results and practical performance. Because of the superiority of SGHMC over SGLD in terms of convergence rate in a wide range of machine learning tasks, it would be a natural question whether such variance reduction techniques can be applied on SGHMC to achieve better results than variance-reduced SGLD.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found