Interpreting Multivariate Interactions in DNNs
Zhang, Hao, Xie, Yichen, Zheng, Longjie, Zhang, Die, Zhang, Quanshi
–arXiv.org Artificial Intelligence
This paper aims to explain deep neural networks (DNNs) from the perspective of multivariate interactions. In this paper, we define and quantify the significance of interactions among multiple input variables of the DNN. Input variables with strong interactions usually form a coalition and reflect prototype features, which are memorized and used by the DNN for inference. We define the significance of interactions based on the Shapley value, which is designed to assign the attribution value of each input variable to the inference. We have conducted experiments with various DNNs. Experimental results have demonstrated the effectiveness of the proposed method. Deep neural networks (DNNs) have exhibited significant success in many tasks, and the interpretability of DNNs has received increasing attention in recent years. Most previous studies of post-hoc explanation of DNNs either explain DNN semantically/visually Lundberg & Lee (2017); Ribeiro et al. (2016), or analyze the representation capacity of DNNs Higgins et al. (2017); Achille & Soatto (2018); Fort et al. (2019); Liang et al. (2019). In this paper, we propose a new perspective to explain a trained DNN, i.e. quantifying interactions among input variables that are used by the DNN during the inference process. Each input variable of a DNN usually does not work individually. Instead, input variables may cooperate with other variables to make inferences. We can consider the strongly interacted input variables to form a prototype feature (or a coalition), which is memorized by the DNN. For example, the face is a prototype feature for person detection, which is comprised of the eyes, nose, and mouth.
arXiv.org Artificial Intelligence
Oct-15-2020
- Genre:
- Research Report > New Finding (0.48)
- Industry:
- Media > Film (1.00)
- Leisure & Entertainment (1.00)
- Technology: