Statistical Learning
Appendix A Proof of Theoretical results
A.1 Proof of Proposition 1 and 3 To prove Proposition 1, we first need the following lemma: Readers may refer to [47] for the proof of this lemma. Let's first consider the left handside, The first inequality is due to information processing inequality. The compactness assumption in Proposition 2 seems restrictive, since BNNs with Gaussian priors on weights will break the compactness assumption. Indeed, the assumptions in proposition 2 are merely sufficient conditions. In this section, we discuss the non-parametric counter part of Proposition 2, i.e., is the grid functional KL between a parametric model and a Gaussian process is still finite?
Functional Variational Inference based on Stochastic Process Generators
Bayesian inference in the space of functions has been an important topic for Bayesian modeling in the past. In this paper, we propose a new solution to this problem called Functional V ariational Inference (FVI). In FVI, we minimize a divergence in function space between the variational distribution and the posterior process.