condresblock
TelescopingDensity-RatioEstimation: SupplementaryMaterial
Dashed boxes denote layers that are not always present. For the MNIST energy-based modelling experiments, we use average pooling operations since other work [19, 4] has found this to produce higher quality samples than max pooling. We note that that another commonly used feature in recent GAN and EBM architectures is Spectral Normalisation (SN) [13]. Our preliminary experiments suggested that SN was not beneficial for performance. As stated in the main text, the number and (in the case of linear combinations) the spacing of the waymarks aretreated ashyperparameters.
Telescoping Density-Ratio Estimation: Supplementary Material
We did not investigate the use of BN, since many energy-based modelling papers (e.g. Our preliminary experiments suggested that SN was not beneficial for performance. As stated in the main text, the number and (in the case of linear combinations) the spacing of the waymarks are treated as hyperparameters. As illustrated by our sensitivity analysis for MNIST (see Figure 5) it seems that, past a certain point, performance plateaus with the addition of extra waymarks. Table 1 shows the grid-searches we performed for all experiments.
Improved Techniques for Training Score-Based Generative Models
Score-based generative models can produce high quality image samples comparable to GANs, without requiring adversarial optimization. However, existing training procedures are limited to images of low resolution (typically below 32 32), and can be unstable under some settings. We provide a new theoretical analysis of learning and sampling from score-based models in high dimensional spaces, explaining existing failure modes and motivating new solutions that generalize across datasets. To enhance stability, we also propose to maintain an exponential moving average of model weights. With these improvements, we can scale scorebased generative models to various image datasets, with diverse resolutions ranging from 64 64 to 256 256. Our score-based models can generate high-fidelity samples that rival best-in-class GANs on various image datasets, including CelebA, FFHQ, and several LSUN categories.