Statistical Learning
Rankmax: An Adaptive Projection Alternative to the Softmax Function
Many machine learning models involve mapping a score vector to a probability vector. Usually, this is done by projecting the score vector onto a probability simplex, and such projections are often characterized as Lipschitz continuous approximations of the argmax function, whose Lipschitz constant is controlled by a parameter that is similar to a softmax temperature.
Supplementary Material to Improving Inference for Neural Compression S1 Stochastic Annealing
Here we provide conceptual illustrations of our stochastic annealing idea on a simple example. As mentioned in Section 3.3 and 4, lossy bits-back modifies the above Base Hyperprior as follows: All methods were tuned on a best-effort basis to ensure convergence, except that STE consistently encountered convergence issues even with a tiny learning rate (see [Yin et al., 2019]). The rate-distortion results for MAP and STE were calculated with early stopping (i.e., using the intermediate Figures in the bottom row focus on the same cropped region of images in the top row. RGB; higher values are better.