Cyclical Annealing Schedule: A Simple Approach to Mitigating KL Vanishing
Fu*, Hao, Li*, Chunyuan, Liu, Xiaodong, Gao, Jianfeng, Celikyilmaz, Asli, Carin, Lawrence
–arXiv.org Artificial Intelligence
One path is conditioned on the latent in many NLP tasks, including language modeling codes, and the other path is conditioned on previously (Bowman et al., 2015; Miao et al., 2016), generated words. KL vanishing happens because dialog response generation (Zhao et al., 2017; (i) the first path can easily get blocked, due Wen et al., 2017), semi-supervised text classification to the lack of good latent codes at the beginning of (Xu et al., 2017), controllable text generation decoder training; (ii) the easiest solution that an (Hu et al., 2017), and text compression (Miao expressive decoder can learn is to ignore the latent and Blunsom, 2016). A prominent component of a code, and relies on the other path only for decoding. VAE is the distribution-based latent representation To remedy this issue, a promising approach is for text sequence observations. This flexible representation to remove the blockage in the first path, and feed allows the VAE to explicitly model holistic meaningful latent codes in training the decoder, so properties of sentences, such as style, topic, and that the decoder can easily adopt them to generate high-level linguistic and semantic features.
arXiv.org Artificial Intelligence
Mar-25-2019