Goto

Collaborating Authors

 Technology




AApproximate Target Maximum Welfare Minimum Relative Entropy Equilbiria We use a Minimum Relative Entropy (RME) (also known as minimum KL divergence) Pa (a)ln

Neural Information Processing Systems

This objective is similar to Maximum Entropy Correlated Equilibrium (MECE) [48], and the proofs here are similar to the framework set out there. A drawback of MECE is that it is not easy to determine the minimum p permissible. If we choose p that does not permit a valid solution, then the parameters will diverge. We can circumvent this problem by optimizing the distance to a target ห† p. And ยตis for balancing the linear objective.




Supplementary Material for Mixture weights optimisation for Alpha-Divergence Variational Inference Kamรฉlia Daudel1,2, Randal Douc3

Neural Information Processing Systems

Assume that p and k are as in (A1). Then, the two following assertions hold. A.3 The case ฮฑ < 1 for the Power Descent algorithm Let ฮฑ = 1, ฮท (0,1], ฮบbe such that (ฮฑ 1)ฮบ 0and let the initial probability measure ยต1 M1(T) be such that ฮจฮฑ(ยต1) < . A common way to approximate intractable integrals of the form (16) is to resort to Importance Sampling methods and in that case we are also interested in ensuring that the support of the variational approximation q Q (with q = ยตk in our case) is included in the support of p. Seeking to solve the Variational Inference optimation problem inf Dฮฑ(ยตK||P) for ฮฑ < 1 enables this to happen, as opposed to the case ฮฑ 1 for which the ฮฑ-divergenve exhibits the so-called mode-seeking property [2, 3, 4]. As a whole, well-chosen samplers and variance reduction methods appear to be a necessity even in the case ฮฑ = 1 so that the obtained Monte Carlo estimator of ฮธ 7 bยต,ฮฑ(ฮธ)do not suffer from a too large variance.


Mixture weights optimisation for Alpha-Divergence Variational Inference

Neural Information Processing Systems

This paper focuses on ฮฑ-divergence minimisation methods for Variational Inference. We consider the case where the posterior density is approximated by a mixture model and we investigate algorithms optimising the mixture weights of this mixture model by ฮฑ-divergence minimisation, without any information on the underlying distribution of its mixture components parameters. The Power Descent, defined for all ฮฑ = 1, is one such algorithm and we establish in our work the full proof of its convergence towards the optimal mixture weights when ฮฑ < 1. Since the ฮฑ-divergence recovers the widely-used exclusive Kullback-Leibler when ฮฑ 1, we then extend the Power Descent to the case ฮฑ = 1 and show that we obtain an Entropic Mirror Descent. This leads us to investigate the link between Power Descent and Entropic Mirror Descent: first-order approximations allow us to introduce the Rรฉnyi Descent, a novel algorithm for which we prove an O(1/N) convergence rate. Lastly, we compare numerically the behavior of the unbiased Power Descent and of the biased Rรฉnyi Descent and we discuss the potential advantages of one algorithm over the other.


Auditing Fairness by Betting

Neural Information Processing Systems

We provide practical, efficient, and nonparametric methods for auditing the fairness of deployed classification and regression models. Whereas previous work relies on a fixed-sample size, our methods are sequential and allow for the continuous monitoring of incoming data, making them highly amenable to tracking the fairness of real-world systems. We also allow the data to be collected by a probabilistic policy as opposed to sampled uniformly from the population. This enables auditing to be conducted on data gathered for another purpose. Moreover, this policy may change over time and different policies may be used on different subpopulations. Finally, our methods can handle distribution shift resulting from either changes to the model or changes in the underlying population. Our approach is based on recent progress in anytime-valid inference and game-theoretic statistics--the "testing by betting" framework in particular. These connections ensure that our methods are interpretable, fast, and easy to implement. We demonstrate the efficacy of our approach on three benchmark fairness datasets.