Review for NeurIPS paper: Gradient Estimation with Stochastic Softmax Tricks

Neural Information Processing Systems 

Summary and Contributions: Update after the author response: I want to thank the authors for clarifying how exactly KL between prior and approximate posterior is calculated in VI set-up. Usually, an "interesting" inductive bias / prior distribution is formulated in the original combinatorial space X rather than utility space U. Hence, I believe it would be beneficial for the potential reader if this limitation is mentioned explicitly in the paper. The paper is concerned with the task of estimating the gradient of the following form: d E_{X p_\theta}[L(X)] / d\theta. Where X represents a combinatorial object (e.g. This loss is ubiquitous in variational inference approach to latent variable models with structured latent variables.