Accuracy
A Proof of Theorem 3.1
As proved by Feng et al. (2021), the binary cross-entropy loss We include more results on teacher model and teacher model + {DRO (Hashimoto et al., 2018) /ARL (Lahoti et al., 2020) / FairRF (Zhao et al., 2022) /our knowledge distillation} in Tab. Effect of our label smoothing can be observed by comparing between "Teacher (with hard label)" and "Teacher (with softmax/linear label)" in the 6 tables. Here the capacity is the same, the only difference is the label smoothing. Here the training method is the same, only difference is capacity. Table 8: Results on COMP AS dataset with sensitive attribute race .
Appendices A Proofs A.1 Proof of Proposition
Here we proved that (1) and (2) are equivalent; (1) and (3) are equivalent. Proposition 3. 14 Lemma 2. Given With the lemma above, we now present the proof of Proposition 3. B.1 Example Implementation We provide an example implementation of Algorithm 2 in Listing 1. 17 1 Based on exponentiation by squaring. Best results are in bold. Based on the observation, Wei et al. Our method identified a different source of gradient vanishing caused by the small coefficients for higher-order terms in DAG constraints.