fast axiomatic attribution
Fast Axiomatic Attribution for Neural Networks
Mitigating the dependence on spurious correlations present in the training dataset is a quickly emerging and important topic of deep learning. Recent approaches include priors on the feature attribution of a deep neural network (DNN) into the training process to reduce the dependence on unwanted features. However, until now one needed to trade off high-quality attributions, satisfying desirable axioms, against the time required to compute them. This in turn either led to long training times or ineffective attribution priors. In this work, we break this trade-off by considering a special class of efficiently axiomatically attributable DNNs for which an axiomatic feature attribution can be computed with only a single forward/backward pass. We formally prove that nonnegatively homogeneous DNNs, here termed $\mathcal{X}$-DNNs, are efficiently axiomatically attributable and show that they can be effortlessly constructed from a wide range of regular DNNs by simply removing the bias term of each layer. Various experiments demonstrate the advantages of $\mathcal{X}$-DNNs, beating state-of-the-art generic attribution methods on regular DNNs for training with attribution priors.
Fast Axiomatic Attribution for Neural Networks - Supplemental Material - Robin Hesse
In the proof of Proposition 3.2, we make use of the property that the derivative of a ( k 1) If the pooling function is linear, homogeneity implicitly holds. If the pooling function is selecting values based on their relative ordering, we consider two cases. For Expected Gradients to satisfy the same axioms that are satisfied by Integrated Gradients, convergence must have occurred, which can only be expected after multiple gradient evaluations. Gradient, Sensitivity (a) is also not satisfied by Expected Gradients in general. Sensitivity (b): As the gradient w.r .t. an irrelevant feature will always be zero, Sensitivity (b) Why is nonnegative homogeneity a desirable axiom for attribution methods?
Fast Axiomatic Attribution for Neural Networks
Mitigating the dependence on spurious correlations present in the training dataset is a quickly emerging and important topic of deep learning. Recent approaches include priors on the feature attribution of a deep neural network (DNN) into the training process to reduce the dependence on unwanted features. However, until now one needed to trade off high-quality attributions, satisfying desirable axioms, against the time required to compute them. This in turn either led to long training times or ineffective attribution priors. In this work, we break this trade-off by considering a special class of efficiently axiomatically attributable DNNs for which an axiomatic feature attribution can be computed with only a single forward/backward pass. We formally prove that nonnegatively homogeneous DNNs, here termed \mathcal{X} -DNNs, are efficiently axiomatically attributable and show that they can be effortlessly constructed from a wide range of regular DNNs by simply removing the bias term of each layer.