Bayesian Inference
Recent Advances in Autoencoder-Based Representation Learning
Tschannen, Michael, Bachem, Olivier, Lucic, Mario
Learning useful representations with little or no supervision is a key challenge in artificial intelligence. We provide an in-depth review of recent advances in representation learning with a focus on autoencoder-based models. To organize these results we make use of meta-priors believed useful for downstream tasks, such as disentanglement and hierarchical organization of features. In particular, we uncover three main mechanisms to enforce such properties, namely (i) regularizing the (approximate or aggregate) posterior distribution, (ii) factorizing the encoding and decoding distribution, or (iii) introducing a structured prior distribution. While there are some promising results, implicit or explicit supervision remains a key enabler and all current methods use strong inductive biases and modeling assumptions. Finally, we provide an analysis of autoencoder-based representation learning through the lens of rate-distortion theory and identify a clear tradeoff between the amount of prior knowledge available about the downstream tasks, and how useful the representation is for this task.
Local Probabilistic Model for Bayesian Classification: a Generalized Local Classification Model
Mao, Chengsheng, Lu, Lijuan, Hu, Bin
In Bayesian classification, it is important to establish a probabilistic model for each class for likelihood estimation. Most of the previous methods modeled the probability distribution in the whole sample space. However, real-world problems are usually too complex to model in the whole sample space; some fundamental assumptions are required to simplify the global model, for example, the class conditional independence assumption for naive Bayesian classification. In this paper, with the insight that the distribution in a local sample space should be simpler than that in the whole sample space, a local probabilistic model established for a local region is expected much simpler and can relax the fundamental assumptions that may not be true in the whole sample space. Based on these advantages we propose establishing local probabilistic models for Bayesian classification. In addition, a Bayesian classifier adopting a local probabilistic model can even be viewed as a generalized local classification model; by tuning the size of the local region and the corresponding local model assumption, a fitting model can be established for a particular classification problem. The experimental results on several real-world datasets demonstrate the effectiveness of local probabilistic models for Bayesian classification.
Surrogate-assisted Bayesian inversion for landscape and basin evolution models
Chandra, Rohitash, Azam, Danial, Kapoor, Arpit, Mรผller, R. Dietmar
The complex and computationally expensive features of the forward landscape and sedimentary basin evolution models pose a major challenge in the development of efficient inference and optimization methods. Bayesian inference provides a methodology for estimation and uncertainty quantification of free model parameters. In our previous work, parallel tempering Bayeslands was developed as a framework for parameter estimation and uncertainty quantification for the landscape and basin evolution modelling software Badlands. Parallel tempering Bayeslands features high-performance computing with dozens of processing cores running in parallel to enhance computational efficiency. Although parallel computing is used, the procedure remains computationally challenging since thousands of samples need to be drawn and evaluated. In large-scale landscape and basin evolution problems, a single model evaluation can take from several minutes to hours, and in certain cases, even days. Surrogate-assisted optimization has been with successfully applied to a number of engineering problems. This motivates its use in optimisation and inference methods suited for complex models in geology and geophysics. Surrogates can speed up parallel tempering Bayeslands by developing computationally inexpensive surrogates to mimic expensive models. In this paper, we present an application of surrogate-assisted parallel tempering where that surrogate mimics a landscape evolution model including erosion, sediment transport and deposition, by estimating the likelihood function that is given by the model. We employ a machine learning model as a surrogate that learns from the samples generated by the parallel tempering algorithm. The results show that the methodology is effective in lowering the overall computational cost significantly while retaining the quality of solutions.
Encoding prior knowledge in the structure of the likelihood
Knollmรผller, Jakob, Enรlin, Torsten A.
The inference of deep hierarchical models is problematic due to strong dependencies between the hierarchies. We investigate a specific transformation of the model parameters based on the multivariate distributional transform. This transformation is a special form of the reparametrization trick, flattens the hierarchy and leads to a standard Gaussian prior on all resulting parameters. The transformation also transfers all the prior information into the structure of the likelihood, hereby decoupling the transformed parameters a priori from each other. A variational Gaussian approximation in this standardized space will be excellent in situations of relatively uninformative data. Additionally, the curvature of the log-posterior is well-conditioned in directions that are weakly constrained by the data, allowing for fast inference in such a scenario. In an example we perform the transformation explicitly for Gaussian process regression with a priori unknown correlation structure. Deep models are inferred rapidly in highly and slowly in poorly informed situations. The flat model show exactly the opposite performance pattern. A synthesis of both, the deep and the flat perspective, provides their combined advantages and overcomes the individual limitations, leading to a faster inference.
Spiking Neural Networks: A Stochastic Signal Processing Perspective
Jang, Hyeryung, Simeone, Osvaldo, Gardner, Brian, Grรผning, Andrรฉ
Spiking Neural Networks (SNNs) are distributed systems whose computing elements, or neurons, are characterized by analog internal dynamics and by digital and sparse inter-neuron, or synaptic, communications. The sparsity of the synaptic spiking inputs and the corresponding event-driven nature of neural processing can be leveraged by hardware implementations to obtain significant energy reductions as compared to conventional Artificial Neural Networks (ANNs). SNNs can be used not only as coprocessors tocarry out given computing tasks, such as classification, but also as learning machines that adapt their internal parameters, e.g., their synaptic weights, on the basis of data and of a learning criterion. This paper provides an overview of models, learning rules, and applications of SNNs from the viewpoint of stochastic signal processing. INTRODUCTION Artificial Neural Networks (ANNs) have become the de-facto standard tool to carry out supervised, unsupervised, and reinforcement learning tasks. Their recent successes range from image classifiers that outperform human experts in medical diagnosis to machines that defeat professional players at complex games such as Go.
Bayesian Spectral Deconvolution Based on Poisson Distribution: Bayesian Measurement and Virtual Measurement Analytics (VMA)
Nagata, Kenji, Mototake, Yoh-ichi, Muraoka, Rei, Sasaki, Takehiko, Okada, Masato
In this paper, we propose a new method of Bayesian measurement for spectral deconvolution, which regresses spectral data into the sum of unimodal basis function such as Gaussian or Lorentzian functions. Bayesian measurement is a framework for considering not only the target physical model but also the measurement model as a probabilistic model, and enables us to estimate the parameter of a physical model with its confidence interval through a Bayesian posterior distribution given a measurement data set. The measurement with Poisson noise is one of the most effective system to apply our proposed method. Since the measurement time is strongly related to the signal-to-noise ratio for the Poisson noise model, Bayesian measurement with Poisson noise model enables us to clarify the relationship between the measurement time and the limit of estimation. In this study, we establish the probabilistic model with Poisson noise for spectral deconvolution. Bayesian measurement enables us to perform virtual and computer simulation for a certain measurement through the established probabilistic model. This property is called "Virtual Measurement Analytics(VMA)" in this paper. We also show that the relationship between the measurement time and the limit of estimation can be extracted by using the proposed method in a simulation of synthetic data and real data for XPS measurement of MoS$_2$.
From Adaptive Kernel Density Estimation to Sparse Mixture Models
Schretter, Colas, Sun, Jianyong, Schelkens, Peter
We introduce a balloon estimator in a generalized expectation-maximization method for estimating all parameters of a Gaussian mixture model given one data sample per mixture component. Instead of limiting explicitly the model size, this regularization strategy yields low-complexity sparse models where the number of effective mixture components reduces with an increase of a smoothing probability parameter $\mathbf{P>0}$. This semi-parametric method bridges from non-parametric adaptive kernel density estimation (KDE) to parametric ordinary least-squares when $\mathbf{P=1}$. Experiments show that simpler sparse mixture models retain the level of details present in the adaptive KDE solution.
Variational Bayesian Complex Network Reconstruction
Xu, Shuang, Zhang, Chun-Xia, Wang, Pei, Zhang, Jiangshe
The networked systems are ubiquitous in many fields, including social-tech science [1, 2], bioinformatics [3-6], epidemic dynamics [7-9] and power grid [10, 11]. However, as is often the case, it is not able to observe the topology of a network, while data generated by this network are available. Therefore, in interdisciplinary science, one of the most important but challenging problems is to reconstruct the complex network from the observed data or time series [12]. This problem has been widely investigated in the past three decades, where the classical method is the delay-coordinate embedding method proposed by Takens [13], which, nevertheless, is only suitable for small-scale networks. Nowadays, with the advent of big data era [14], it is of great urgency solve this issue for large-scale complex networks. Suppose that a complex network consists of N nodes, in practice we are often given the time series of the states for the N nodes. Generally speaking, the core idea of many data-driven network reconstruction investigations is to first calculate the correlation between two nodes. Then, a threshold can be set mutually or automatically to make the network binary.
Finding dissimilar explanations in Bayesian networks: Complexity results
Finding the most probable explanation for observed variables in a Bayesian network is a notoriously intractable problem, particularly if there are hidden variables in the network. In this paper we examine the complexity of a related problem, that is, the problem of finding a set of sufficiently dissimilar, yet all plausible, explanations. Applications of this problem are, e.g., in search query results (you won't want 10 results that all link to the same website) or in decision support systems. We show that the problem of finding a 'good enough' explanation that differs in structure from the best explanation is at least as hard as finding the best explanation itself.
Model-Based Learning of Turbulent Flows using Mobile Robots
Khodayi-mehr, Reza, Zavlanos, Michael M.
Abstract--In this paper we consider the problem of modelbased learningof turbulent flows using mobile robots. The key idea is to use empirical data to improve on numerical estimates of time-averaged flow properties that can be obtained using Reynolds-Averaged Navier Stokes (RANS) models. RANS models are computationally efficient and provide global knowledge of the flow but they also rely on simplifying assumptions and require experimental validation. In this paper, we instead construct statistical models of the flow properties using Gaussian Processes (GPs) and rely on the numerical solutions obtained from RANS models to inform their mean. We then utilize Bayesian inference to incorporate empirical measurements of the flow into these GPs, specifically, measurements of the time-averaged velocity and turbulent intensity fields. Moreover, it accounts for measurement noise by systematically incorporating it in the GP models. To obtain the velocity and turbulent intensity measurements, we design a cost-effective mobile robot sensor that collects and analyzes instantaneous velocity readings. We control this mobile robot through a sequence of waypoints that maximize the information content of the corresponding measurements. The end result is a posterior distribution of the flow field that better approximates the real flow and also quantifies the uncertainty in the flow properties. We present experimental results that demonstrate considerable improvement in the prediction of the flow properties compared to pure numerical simulations. I. INTRODUCTION Knowledge of turbulent flow properties, e.g., velocity and turbulent intensity, is of paramount importance for many engineering applications.At larger scales, these properties are used for the study of ocean currents and their effects on aquatic life, [1], [2], [3], meteorology, [4], bathymetry, [5], and localization of atmospheric pollutants, [6], to name a few. At smaller scales, knowledge of flow fields is important in applications ranging from optimal HVAC of residential buildings for human comfort, [7], to design of drag-efficient bodies in aerospace and automotive industries, [8]. At even smaller scales, the characteristics of velocity fluctuations in vessels are important for vascular pathology and diagnosis, [9] or for the control of bacteria-inspired uniflagellar robots, [10]. Another important application that requires global knowledge of the velocity field is chemical source identification in advection-diffusion transport systems, [11], [12], [13].