Goto

Collaborating Authors

 base loss


On Group Sufficiency Under Label Bias

Neural Information Processing Systems

Real-world classification datasets often contain label bias, where observed labels differ systematically from the true labels at different rates for different demographic groups. Machine learning models trained on such datasets may then exhibit disparities in predictive performance across these groups. In this work, we characterize the problem of learning fair classification models with respect to the underlying ground truth labels when given only label biased data. We focus on the particular fairness definition of group sufficiency, i.e. equal calibration of risk scores across protected groups. We theoretically show that enforcing fairness with respect to label biased data necessarily results in group miscalibration with respect to the true labels. We then propose a regularizer which minimizes an upper bound on the sufficiency gap by penalizing a conditional mutual information term. Across experiments on eight tabular, image, and text datasets with both synthetic and real label noise, we find that our method reduces the sufficiency gap by up to 7.2% with no significant decrease in overall accuracy.



raised by multiple reviewers and next respond to individual questions

Neural Information Processing Systems

We thank all the reviewers for their feedback and pointers to relevant papers. This includes (Kendall et al., 2018), where they learn Kendall et al. 2018), we consider different loss functions on the same output space. There are specific reasons we did not use several multi-task learning algorithms mentioned by REV4 as baselines. Kendall et al. (2018) assumes that all base losses are applications of the same function (max likelihood in this case) We don't see how this method can be extended to our scenario where base losses do not necessarily Moreover, our regularization admits a very different nature. However, directly normalizing the base losses was sufficient for our experiments.


Flexible risk design using bi-directional dispersion

arXiv.org Artificial Intelligence

Many novel notions of "risk" (e.g., CVaR, tilted risk, DRO risk) have been proposed and studied, but these risks are all at least as sensitive as the mean to loss tails on the upside, and tend to ignore deviations on the downside. We study a complementary new risk class that penalizes loss deviations in a bi-directional manner, while having more flexibility in terms of tail sensitivity than is offered by mean-variance. This class lets us derive high-probability learning guarantees without explicit gradient clipping, and empirical tests using both simulated and real data illustrate a high degree of control over key properties of the test loss distribution incurred by gradient-based learners.


Student-Teacher Learning from Clean Inputs to Noisy Inputs

arXiv.org Machine Learning

Feature-based student-teacher learning, a training method that encourages the student's hidden features to mimic those of the teacher network, is empirically successful in transferring the knowledge from a pre-trained teacher network to the student network. Furthermore, recent empirical results demonstrate that, the teacher's features can boost the student network's generalization even when the student's input sample is corrupted by noise. However, there is a lack of theoretical insights into why and when this method of transferring knowledge can be successful between such heterogeneous tasks. We analyze this method theoretically using deep linear networks, and experimentally using nonlinear networks. We identify three vital factors to the success of the method: (1) whether the student is trained to zero training loss; (2) how knowledgeable the teacher is on the clean-input problem; (3) how the teacher decomposes its knowledge in its hidden features. Lack of proper control in any of the three factors leads to failure of the student-teacher learning method.