Goto

Collaborating Authors

 Bayesian Inference


Learning Directed Graphical Models with Optimal Transport

arXiv.org Artificial Intelligence

Estimating the parameters of a probabilistic directed graphical model from incomplete data remains a long-standing challenge. This is because, in the presence of latent variables, both the likelihood function and posterior distribution are intractable without further assumptions about structural dependencies or model classes. While existing learning methods are fundamentally based on likelihood maximization, here we offer a new view of the parameter learning problem through the lens of optimal transport. This perspective licenses a general framework that operates on any directed graphs without making unrealistic assumptions on the posterior over the latent variables or resorting to black-box variational approximations. We develop a theoretical framework and support it with extensive empirical evidence demonstrating the flexibility and versatility of our approach. Across experiments, we show that not only can our method recover the ground-truth parameters but it also performs comparably or better on downstream applications, notably the non-trivial task of discrete representation learning.


Compositional Generative Modeling: A Single Model is Not All You Need

arXiv.org Artificial Intelligence

Large monolithic generative models trained on massive amounts of data have become an increasingly dominant approach in AI research. In this paper, we argue that we should instead construct large generative systems by composing smaller generative models together. We show how such a compositional generative approach enables us to learn distributions in a more data-efficient manner, enabling generalization to parts of the data distribution unseen at training time. We further show how this enables us to program and construct new generative models for tasks completely unseen at training. Finally, we show that in many cases, we can discover separate compositional components from data.


Bayesian Causal Inference with Gaussian Process Networks

arXiv.org Artificial Intelligence

Quantifying the causal relationships from purely observational data between variables in a system is a problem that has attracted great attention in the fields of statistics and machine learning. Full knowledge of the causal relations allows predicting the outcome of direct manipulations on the system, which can generally only be known from interventional data obtained by performing experiments such as randomized controlled trials (Eberhardt and Scheines, 2007). Predicting the effect of such manipulations without the need of costly or infeasible experiments is of great practical relevance, specifically in the fields of computational biology (Sachs et al., 2005), medicine (Richens et al., 2020) or AI (Schölkopf, 2022), since a central question concerns how a complex system will react to some treatment or outside influence of the user. Pearl's rules of do-calculus (Pearl, 2000) allow computing the intervention distributions resulting from these external manipulations from the joint distribution of the set of random variables together with a Directed Acyclic Graph (DAG). The DAG represents the qualitative causal relationships among the variables; each node in the graph represents a variable and a directed edge indicates a direct causal effect. Probabilistic models that are based on such DAGs, commonly called causal Bayesian Networks (BNs), provide conventional grounds for probabilistic causal inference, due to their compact representation of the joint distribution and their intuitive graphical description of the causal structure.


Adaptive Crowdsourcing Via Self-Supervised Learning

arXiv.org Artificial Intelligence

Common crowdsourcing systems average estimates of a latent quantity of interest provided by many crowdworkers to produce a group estimate. We develop a new approach -- predict-each-worker -- that leverages self-supervised learning and a novel aggregation scheme. This approach adapts weights assigned to crowdworkers based on estimates they provided for previous quantities. When skills vary across crowdworkers or their estimates correlate, the weighted sum offers a more accurate group estimate than the average. Existing algorithms such as expectation maximization can, at least in principle, produce similarly accurate group estimates. However, their computational requirements become onerous when complex models, such as neural networks, are required to express relationships among crowdworkers. Predict-each-worker accommodates such complexity as well as many other practical challenges. We analyze the efficacy of predict-each-worker through theoretical and computational studies. Among other things, we establish asymptotic optimality as the number of engagements per crowdworker grows.


Comprehensive Exploration of Synthetic Data Generation: A Survey

arXiv.org Artificial Intelligence

Recent years have witnessed a surge in the popularity of Machine Learning (ML), applied across diverse domains. However, progress is impeded by the scarcity of training data due to expensive acquisition and privacy legislation. Synthetic data emerges as a solution, but the abundance of released models and limited overview literature pose challenges for decision-making. This work surveys 417 Synthetic Data Generation (SDG) models over the last decade, providing a comprehensive overview of model types, functionality, and improvements. Common attributes are identified, leading to a classification and trend analysis. The findings reveal increased model performance and complexity, with neural network-based approaches prevailing, except for privacy-preserving data generation. Computer vision dominates, with GANs as primary generative models, while diffusion models, transformers, and RNNs compete. Implications from our performance evaluation highlight the scarcity of common metrics and datasets, making comparisons challenging. Additionally, the neglect of training and computational costs in literature necessitates attention in future research. This work serves as a guide for SDG model selection and identifies crucial areas for future exploration.


EMO: Earth Mover Distance Optimization for Auto-Regressive Language Modeling

arXiv.org Artificial Intelligence

Neural language models are probabilistic models of human text. They are predominantly trained using maximum likelihood estimation (MLE), which is equivalent to minimizing the forward cross-entropy between the empirical data distribution and the model distribution. However, various degeneration phenomena are still widely observed when decoding from the distributions learned by such models. We establish that the forward cross-entropy is suboptimal as a distance metric for aligning human and model distribution due to its (1) recall-prioritization (2) negative diversity ignorance and (3) train-test mismatch. In this paper, we propose Earth Mover Distance Optimization (EMO) for auto-regressive language modeling. EMO capitalizes on the inherent properties of earth mover distance to address the aforementioned challenges. Due to the high complexity of direct computation, we further introduce a feasible upper bound for EMO to ease end-to-end training. Upon extensive evaluation of language models trained using EMO and MLE. We find that EMO demonstrates a consistently better language modeling performance than MLE across domains. Moreover, EMO demonstrates noteworthy enhancements in downstream performance with minimal fine-tuning on merely 25,000 sentences. This highlights the tremendous potential of EMO as a lightweight calibration method for enhancing large-scale pre-trained language models.


Convergence of Expectation-Maximization Algorithm with Mixed-Integer Optimization

arXiv.org Machine Learning

The convergence of expectation-maximization (EM)-based algorithms typically requires continuity of the likelihood function with respect to all the unknown parameters (optimization variables). The requirement is not met when parameters comprise both discrete and continuous variables, making the convergence analysis nontrivial. This paper introduces a set of conditions that ensure the convergence of a specific class of EM algorithms that estimate a mixture of discrete and continuous parameters. Our results offer a new analysis technique for iterative algorithms that solve mixed-integer non-linear optimization problems. As a concrete example, we prove the convergence of the EM-based sparse Bayesian learning algorithm in [1] that estimates the state of a linear dynamical system with jointly sparse inputs and bursty missing observations. Our results establish that the algorithm in [1] converges to the set of stationary points of the maximum likelihood cost with respect to the continuous optimization variables.


AlphaRank: An Artificial Intelligence Approach for Ranking and Selection Problems

arXiv.org Artificial Intelligence

We introduce AlphaRank, an artificial intelligence approach to address the fixed-budget ranking and selection (R&S) problems. We formulate the sequential sampling decision as a Markov decision process and propose a Monte Carlo simulation-based rollout policy that utilizes classic R&S procedures as base policies for efficiently learning the value function of stochastic dynamic programming. We accelerate online sample-allocation by using deep reinforcement learning to pre-train a neural network model offline based on a given prior. We also propose a parallelizable computing framework for large-scale problems, effectively combining "divide and conquer" and "recursion" for enhanced scalability and efficiency. Numerical experiments demonstrate that the performance of AlphaRank is significantly improved over the base policies, which could be attributed to AlphaRank's superior capability on the trade-off among mean, variance, and induced correlation overlooked by many existing policies.


Variable selection for Na\"ive Bayes classification

arXiv.org Artificial Intelligence

The Na\"ive Bayes has proven to be a tractable and efficient method for classification in multivariate analysis. However, features are usually correlated, a fact that violates the Na\"ive Bayes' assumption of conditional independence, and may deteriorate the method's performance. Moreover, datasets are often characterized by a large number of features, which may complicate the interpretation of the results as well as slow down the method's execution. In this paper we propose a sparse version of the Na\"ive Bayes classifier that is characterized by three properties. First, the sparsity is achieved taking into account the correlation structure of the covariates. Second, different performance measures can be used to guide the selection of features. Third, performance constraints on groups of higher interest can be included. Our proposal leads to a smart search, which yields competitive running times, whereas the flexibility in terms of performance measure for classification is integrated. Our findings show that, when compared against well-referenced feature selection approaches, the proposed sparse Na\"ive Bayes obtains competitive results regarding accuracy, sparsity and running times for balanced datasets. In the case of datasets with unbalanced (or with different importance) classes, a better compromise between classification rates for the different classes is achieved.


Algorithmic Robust Forecast Aggregation

arXiv.org Artificial Intelligence

Forecast aggregation combines the predictions of multiple agents into a more accurate prediction. With forecast aggregation, decision-makers can reduce error, diversify risk and enhance accuracy based on the collective knowledge of agents compared to any single agent, thereby advancing the common good. Forecast aggregation is commonly used in many domains to generate more informed predictions for various variables, such as weather in weather forecasting, the spread of infectious diseases in public health, the outcome of games in sports, fuel prices in energy, and GDP growth in economics. In practice, one crucial challenge of forecast aggregation is that the aggregator may not have full knowledge of the information structure and the agents. Without this prior knowledge, the aggregator cannot employ Bayes rules to combine the forecasts optimally. Traditional prior-free aggregation methods, such as simple averaging, are especially bad on some information structures. For example, in weather forecasting, assume the prior probability of raining tomorrow is 30%, and there are two agents who will receive a conditionally independent binary signal (Low or High). Agents will report their posterior, which is 10% given the Low signal and 50% given the High signal. When both agents report 50%, the simple averaging will also output 50%.