Masked Generative Modeling with Enhanced Sampling Scheme

Lee, Daesoo, Aune, Erlend, Malacarne, Sara

arXiv.org Machine Learning 

In recent years, generative modeling has made significant advancements. The mainstream frameworks for generative image modeling have evolved through various phases, initially utilizing Variational AutoEncoder (VAE) [1], then progressing to Generative Adversarial Network (GAN) [2], and eventually Vector Quantized-Variational AutoEncoder (VQ-VAE) [3] and diffusion models [4]. Masked generative modeling on VQ-tokens and diffusion models have demonstrated state-of-the-art (SOTA) performance in the last couple of years on image modeling, audio modeling, and time series modeling [5, 6, 7, 8, 9, 10]. Both masked generative modeling and diffusion models utilize sampling schemes to iteratively unmask/predict tokens or denoise noisy inputs. In recent years, diffusion models' sampling schemes have garnered considerable attention [11, 12]. While the sampling scheme for masked generative modeling has not received the same level of attention, notable progress has been made in [13, 14, 10]. This paper identifies certain limitations of the masked generative modeling sampling schemes in [13, 14, 10], and introduces a novel and improved sampling scheme which we call Enhanced Sampling Scheme (ESS) to address these limitations. ESS comprises three stages: 1) Iterative non-autoregressive decoding, as proposed in [13], 2) Critical Reverse Sampling, and 3) Critical Resampling using self-Token-Critic. Figure 1 provides an overview of the method, and Sect. 3 presents a detailed explanation of ESS.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found