Goto

Collaborating Authors

 Bayesian Learning


Reliable Uncertain Evidence Modeling in Bayesian Networks by Credal Networks

arXiv.org Artificial Intelligence

A reliable modeling of uncertain evidence in Bayesian networks based on a set-valued quantification is proposed. Both soft and virtual evidences are considered. We show that evidence propagation in this setup can be reduced to standard updating in an augmented credal network, equivalent to a set of consistent Bayesian networks. A characterization of the computational complexity for this task is derived together with an efficient exact procedure for a subclass of instances. In the case of multiple uncertain evidences over the same variable, the proposed procedure can provide a set-valued version of the geometric approach to opinion pooling.


Designing Random Graph Models Using Variational Autoencoders With Applications to Chemical Design

arXiv.org Machine Learning

From left to right, given a graph G with a set of node features F and edge weights Y, the encoder aggregates information from a different number of hops j K away for each nodev G into an embedding vectorc v(j). To do so, it uses a feedforward network to propagate information between different search depths, which is parametrized by a set of weight matrices W j . This embedding vectors are then fed into a differentiable functionฯ† enc, which sets the parameters,ยต k andฯƒ k, of several multidimensional Gaussian distributionsq ฯ†, from where the latent representation of each node in the input graph are sampled from. Variational autoencoders are characterized by a probabilistic generative modelp ฮธ(x z) of the observed variablesx R N given the latent variablesz R M, a prior distribution over the latent variablesp(z) and an approximate probabilistic inference modelq ฯ† (z x). In this characterization,p ฮธ and q ฯ† are arbitrary distributions parametrized by two (deep) neural networksฮธ and ฯ† and one can think of the generative model as a probabilistic decoder, which decodes latent variables into observed variables, and the inference model as a probabilistic encoder, which encodes observed variables into latent variables. Ideally, if we use the maximum likelihood principle to train a variational autoencoder, we should optimize the marginal log-likelihood of the observed data, i.e., E D [log p ฮธ(x)], wherep D is the data distribution. Unfortunately, computing logp ฮธ(x) requires marginalization with respect to the the latent variablez, which is typically intractable. Therefore, one resorts to maximizing a variational lower bound or evidence lower bound (ELBO) of the log-likelihood the observed data, i.e., max ฮธ max ฯ† E D [ KL(q ฯ† (z x) p(z)) E q ฯ† (z x)log p ฮธ(x z)] . Finally, note that the quality of this variational lower bound (and thus the quality of the resulting V AE) depends on the expressive ability of the approximate inference modelq ฯ† (z x), which is typically assumed to be a normal distribution whose mean and variance are parametrized by a (deep) neural networkฯ† with the observed datax as an input.


Vertex nomination: The canonical sampling and the extended spectral nomination schemes

arXiv.org Machine Learning

Suppose that one particular block in a stochastic block model is deemed "interesting," but block labels are only observed for a few of the vertices. Utilizing a graph realized from the model, the vertex nomination task is to order the vertices with unobserved block labels into a "nomination list" with the goal of having an abundance of interesting vertices near the top of the list. In this paper we extend and enhance two basic vertex nomination schemes; the canonical nomination scheme ${\mathcal L}^C$ and the spectral partitioning nomination scheme ${\mathcal L}^P$. The canonical nomination scheme ${\mathcal L}^C$ is provably optimal, but is computationally intractable, being impractical to implement even on modestly sized graphs. With this in mind, we introduce a scalable, Markov chain Monte Carlo-based nomination scheme, called the {\it canonical sampling nomination scheme} ${\mathcal L}^{CS}$, that converges to the canonical nomination scheme ${\mathcal L}^{C}$ as the amount of sampling goes to infinity. We also introduce a novel spectral partitioning nomination scheme called the {\it extended spectral partitioning nomination scheme} ${\mathcal L}^{EP}$. Real-data and simulation experiments are employed to illustrate the effectiveness of these vertex nomination schemes, as well as their empirical computational complexity.


Superposition-Assisted Stochastic Optimization for Hawkes Processes

arXiv.org Machine Learning

We consider the learning of multi-agent Hawkes processes, a model containing multiple Hawkes processes with shared endogenous impact functions and different exogenous intensities. In the framework of stochastic maximum likelihood estimation, we explore the associated risk bound. Further, we consider the superposition of Hawkes processes within the model, and demonstrate that under certain conditions such an operation is beneficial for tightening the risk bound. Accordingly, we propose a stochastic optimization algorithm assisted with a diversity-driven superposition strategy, achieving better learning results with improved convergence properties. The effectiveness of the proposed method is verified on synthetic data, and its potential to solve the cold-start problem of sequential recommendation systems is demonstrated on real-world data.


Inferring Tweedie Compound Poisson Mixed Models with Adversarial Variational Bayes

arXiv.org Machine Learning

The Tweedie Compound Poisson-Gamma model is routinely used for modelling non-negative continuous data with a discrete probability mass at zero. Mixed models with random effects account for the covariance structure related to the grouping hierarchy in the data. An important application of Tweedie mixed models is estimating the aggregated loss for insurance policies. However, the intractable likelihood function, the unknown variance function, and the hierarchical structure of mixed effects have presented considerable challenges for drawing inferences on Tweedie. In this study, we tackle the Bayesian Tweedie mixed-effects models via variational approaches. In particular, we empower the posterior approximation by implicit models trained in an adversarial setting. To reduce the variance of gradients, we reparameterize random effects, and integrate out one local latent variable of Tweedie. We also employ a flexible hyper prior to ensure the richness of the approximation. Our method is evaluated on both simulated and real-world data. Results show that the proposed method has smaller estimation bias on the random effects compared to traditional inference methods including MCMC; it also achieves a state-of-the-art predictive performance, meanwhile offering a richer estimation of the variance function.


Maximizing the information learned from finite data selects a simple model

arXiv.org Machine Learning

We use the language of uninformative Bayesian prior choice to study the selection of appropriately simple effective models. We advocate for the prior which maximizes the mutual information between parameters and predictions, learning as much as possible from limited data. When many parameters are poorly constrained by the available data, we find that this prior puts weight only on boundaries of the parameter manifold. Thus it selects a lower-dimensional effective theory in a principled way, ignoring irrelevant parameter directions. In the limit where there is sufficient data to tightly constrain any number of parameters, this reduces to Jeffreys prior. But we argue that this limit is pathological when applied to the hyper-ribbon parameter manifolds generic in science, because it leads to dramatic dependence on effects invisible to experiment.


Extreme Stochastic Variational Inference: Distributed and Asynchronous

arXiv.org Machine Learning

We propose extreme stochastic variational inference (ESVI), an asynchronous and lock-free algorithm to perform variational inference on massive real world datasets. Stochastic variational inference (SVI), the state-of-the-art algorithm for scaling variational inference to large-datasets, is inherently serial. Moreover, it requires the parameters to fit in the memory of a single processor; this is problematic when the number of parameters is in billions. ESVI overcomes these limitations by requiring that each processor only access a subset of the data and a subset of the parameters, thus providing data and model parallelism simultaneously. We demonstrate the effectiveness of ESVI by running Latent Dirichlet Allocation (LDA) on UMBC-3B, a dataset that has a vocabulary of 3 million and a token size of 3 billion. To best of our knowledge, this is an order of magnitude larger than the largest dataset on which results using variational inference have been reported in literature. In our experiments, we found that ESVI outperforms VI and SVI, and also achieves a better quality solution. In addition, we propose a strategy to speed up computation and save memory when fitting large number of topics.


A Bayesian Perspective on Generalization and Stochastic Gradient Descent

arXiv.org Artificial Intelligence

We consider two questions at the heart of machine learning; how can we predict if a minimum will generalize to the test set, and why does stochastic gradient descent find minima that generalize well? Our work responds to Zhang et al. (2016), who showed deep neural networks can easily memorize randomly labeled training data, despite generalizing well on real labels of the same inputs. We show that the same phenomenon occurs in small linear models. These observations are explained by the Bayesian evidence, which penalizes sharp minima but is invariant to model parameterization. We also demonstrate that, when one holds the learning rate fixed, there is an optimum batch size which maximizes the test set accuracy. We propose that the noise introduced by small mini-batches drives the parameters towards minima whose evidence is large. Interpreting stochastic gradient descent as a stochastic differential equation, we identify the "noise scale" $g = \epsilon (\frac{N}{B} - 1) \approx \epsilon N/B$, where $\epsilon$ is the learning rate, $N$ the training set size and $B$ the batch size. Consequently the optimum batch size is proportional to both the learning rate and the size of the training set, $B_{opt} \propto \epsilon N$. We verify these predictions empirically.


Continuous-Time Flows for Efficient Inference and Density Estimation

arXiv.org Machine Learning

Two fundamental problems in unsupervised learning are efficient inference for latent-variable models and robust density estimation based on large amounts of unlabeled data. Algorithms for the two tasks, such as normalizing flows and generative adversarial networks (GANs), are often developed independently. In this paper, we propose the concept of {\em continuous-time flows} (CTFs), a family of diffusion-based methods that are able to asymptotically approach a target distribution. Distinct from normalizing flows and GANs, CTFs can be adopted to achieve the above two goals in one framework, with theoretical guarantees. Our framework includes distilling knowledge from a CTF for efficient inference, and learning an explicit energy-based distribution with CTFs for density estimation. Both tasks rely on a new technique for distribution matching within amortized learning. Experiments on various tasks demonstrate promising performance of the proposed CTF framework, compared to related techniques.


An Improved Bayesian Framework for Quadrature of Constrained Integrands

arXiv.org Machine Learning

Quadrature is the problem of estimating intractable integrals, a problem that arises in many Bayesian machine learning settings. We present an improved Bayesian framework for estimating intractable integrals of specific kinds of constrained integrands. We derive the necessary approximation scheme for a specific and especially useful instantiation of this framework: the use of a log transformation to model non-negative integrands. We also propose a novel method for optimizing the hyperparameters associated with this framework; we optimize the hyperparameters in the original space of the integrand as opposed to in the transformed space, resulting in a model that better explains the actual data. Experiments on both synthetic and real-world data demonstrate that the proposed framework achieves more-accurate estimates using less wall-clock time than previously preposed Bayesian quadrature procedures for non-negative integrands.