Bayesian Inference
Robust Bayesian Subspace Identification for Small Data Sets
Model estimates obtained from traditional subspace identification methods may be subject to significant variance. This elevated variance is aggravated in the cases of large models or of a limited sample size. Common solutions to reduce the effect of variance are regularized estimators, shrinkage estimators and Bayesian estimation. In the current work we investigate the latter two solutions, which have not yet been applied to subspace identification. Our experimental results show that our proposed estimators may reduce the estimation risk up to $40\%$ of that of traditional subspace methods.
Functional Integrative Bayesian Analysis of High-dimensional Multiplatform Genomic Data
Bhattacharyya, Rupam, Henderson, Nicholas, Baladandayuthapani, Veerabhadran
Rapid advancements in collection, processing, and dissemination of multi-platform molecular and genomics (multi-omics, in short) data has resulted in enormous opportunities to aggregate such data in order to understand, prevent, and treat diseases. This has catalyzed development of integrative methods that can collectively mine multiple types and scales of multi-omics data, in order to provide a more holistic view of human disease evolution and progression (Subramanian et al. 2020). Specifically, in the context of cancer, a disease driven predominantly by agglomerations of several molecular changes (Sun et al. 2021), the importance of synthesizing information from multi-platform omics and clinical sources to understand the cellular basis of the disease is even further underscored. Cellular oncological mechanisms, triggered at different molecular levels of the DNA RNA Protein path, can confer profound phenotypic advantages/disadvantages. While significant improvements have been made in multi-omics data integration methods to unveil such mechanisms, focused on both prognosis (Duan et al. 2021) and treatment (Finotello et al. 2020), the precise functions governing them need detailed and data-driven de-novo evaluations. Our work, in the same vein, aims at two different but inter-related scientific axes: (i) selection of biomarkers associated with cancer prognosis and clinical outcomes, and (ii) learning the mechanism of these biomarkers' effects upon such outcomes via integrating upstream molecular information - we provide some additional scientific context below. Classes of Integrative Omics Models First, we briefly discuss existing integrative omics approaches in order to contextualize the need for our framework. Broadly, most of the existing integrative statistical methods can be classified into two categories - horizontal (meta-analysis type) and vertical (multi-omics) integration procedures (Tseng et al. 2015).
A Data-Adaptive Prior for Bayesian Learning of Kernels in Operators
Chada, Neil K., Lang, Quanjun, Lu, Fei, Wang, Xiong
Kernels are efficient in representing nonlocal dependence and they are widely used to design operators between function spaces. Thus, learning kernels in operators from data is an inverse problem of general interest. Due to the nonlocal dependence, the inverse problem can be severely ill-posed with a data-dependent singular inversion operator. The Bayesian approach overcomes the ill-posedness through a non-degenerate prior. However, a fixed non-degenerate prior leads to a divergent posterior mean when the observation noise becomes small, if the data induces a perturbation in the eigenspace of zero eigenvalues of the inversion operator. We introduce a data-adaptive prior to achieve a stable posterior whose mean always has a small noise limit. The data-adaptive prior's covariance is the inversion operator with a hyper-parameter selected adaptive to data by the L-curve method. Furthermore, we provide a detailed analysis on the computational practice of the data-adaptive prior, and demonstrate it on Toeplitz matrices and integral operators. Numerical tests show that a fixed prior can lead to a divergent posterior mean in the presence of any of the four types of errors: discretization error, model error, partial observation and wrong noise assumption. In contrast, the data-adaptive prior always attains posterior means with small noise limits.
Fast and fully-automated histograms for large-scale data sets
Mendizรกbal, Valentina Zelaya, Boullรฉ, Marc, Rossi, Fabrice
G-Enum histograms are a new fast and fully automated method for irregular histogram construction. By framing histogram construction as a density estimation problem and its automation as a model selection task, these histograms leverage the Minimum Description Length principle (MDL) to derive two different model selection criteria. Several proven theoretical results about these criteria give insights about their asymptotic behavior and are used to speed up their optimisation. These insights, combined to a greedy search heuristic, are used to construct histograms in linearithmic time rather than the polynomial time incurred by previous works. The capabilities of the proposed MDL density estimation method are illustrated with reference to other fully automated methods in the literature, both on synthetic and large real-world data sets.
Learning Energy-Based Models With Adversarial Training
Yin, Xuwang, Li, Shiying, Rohde, Gustavo K.
We study a new approach to learning energy-based models (EBMs) based on adversarial training (AT). We show that (binary) AT learns a special kind of energy function that models the support of the data distribution, and the learning process is closely related to MCMC-based maximum likelihood learning of EBMs. We further propose improved techniques for generative modeling with AT, and demonstrate that this new approach is capable of generating diverse and realistic images. Aside from having competitive image generation performance to explicit EBMs, the studied approach is stable to train, is well-suited for image translation tasks, and exhibits strong out-of-distribution adversarial robustness. Our results demonstrate the viability of the AT approach to generative modeling, suggesting that AT is a competitive alternative approach to learning EBMs.
Teaching - CS 221
CS 221 โ Artificial Intelligence My twin brother Afshine and I created this set of illustrated Artificial Intelligence cheatsheets covering the content of the CS 221 class, which I TA-ed in Spring 2019 at Stanford. They can (hopefully!) be useful to all future students of this course as well as to anyone else interested in Artificial Intelligence. You can help us translating them on GitHub!
Probabilistic quantile factor analysis
Korobilis, Dimitris, Schrรถder, Maximilian
This paper extends quantile factor analysis to a probabilistic variant that incorporates regularization and computationally efficient variational approximations. By means of synthetic and real data experiments it is established that the proposed estimator can achieve, in many cases, better accuracy than a recently proposed loss-based estimator. We contribute to the literature on measuring uncertainty by extracting new indexes of low, medium and high economic policy uncertainty, using the probabilistic quantile factor methodology. Medium and high indexes have clear contractionary effects, while the low index is benign for the economy, showing that not all manifestations of uncertainty are the same.
Learning from Heterogeneous Data Based on Social Interactions over Graphs
Bordignon, Virginia, Vlaski, Stefan, Matta, Vincenzo, Sayed, Ali H.
This work proposes a decentralized architecture, where individual agents aim at solving a classification problem while observing streaming features of different dimensions and arising from possibly different distributions. In the context of social learning, several useful strategies have been developed, which solve decision making problems through local cooperation across distributed agents and allow them to learn from streaming data. However, traditional social learning strategies rely on the fundamental assumption that each agent has significant prior knowledge of the underlying distribution of the observations. In this work we overcome this issue by introducing a machine learning framework that exploits social interactions over a graph, leading to a fully data-driven solution to the distributed classification problem. In the proposed social machine learning (SML) strategy, two phases are present: in the training phase, classifiers are independently trained to generate a belief over a set of hypotheses using a finite number of training samples; in the prediction phase, classifiers evaluate streaming unlabeled observations and share their instantaneous beliefs with neighboring classifiers. We show that the SML strategy enables the agents to learn consistently under this highly-heterogeneous setting and allows the network to continue learning even during the prediction phase when it is deciding on unlabeled samples. The prediction decisions are used to continually improve performance thereafter in a manner that is markedly different from most existing static classification schemes where, following training, the decisions on unlabeled data are not re-used to improve future performance.
Statistical Distance Based Deterministic Offspring Selection in SMC Methods
Kviman, Oskar, Koptagel, Hazal, Melin, Harald, Lagergren, Jens
Over the years, sequential Monte Carlo (SMC) and, equivalently, particle filter (PF) theory has gained substantial attention from researchers. However, the performance of the resampling methodology, also known as offspring selection, has not advanced recently. We propose two deterministic offspring selection methods, which strive to minimize the Kullback-Leibler (KL) divergence and the total variation (TV) distance, respectively, between the particle distribution prior and subsequent to the offspring selection. By reducing the statistical distance between the selected offspring and the joint distribution, we obtain a heuristic search procedure that performs superior to a maximum likelihood search in precisely those contexts where the latter performs better than an SMC. For SMC and particle Markov chain Monte Carlo (pMCMC), our proposed offspring selection methods always outperform or compare favorably with the two state-of-the-art resampling schemes on two models commonly used as benchmarks from the literature.
Deep Causal Learning for Robotic Intelligence
This invited review discusses causal learning in the context of robotic intelligence. The paper introduced the psychological findings on causal learning in human cognition, then it introduced the traditional statistical solutions on causal discovery and causal inference. The paper reviewed recent deep causal learning algorithms with a focus on their architectures and the benefits of using deep nets and discussed the gap between deep causal learning and the needs of robotic intelligence.