Who Explains the Explanation? Quantitatively Assessing Feature Attribution Methods
Arias-Duart, Anna, Parés, Ferran, Garcia-Gasulla, Dario
–arXiv.org Artificial Intelligence
Explainability has become a major topic of research in Artificial Intelligence (AI), aimed at increasing trust in models such as Deep Learning (DL) networks. However, trustworthy models cannot be achieved with explainable AI (XAI) methods unless the XAI methods themselves can be trusted. This necessity gave rise to the assessment of XAI methods. To evaluate XAI methods one may assess interpretability, a qualitative measure of how understandable an explanation is to humans Gilpin et al. [2018]. While this is important to guarantee the proper interaction between humans and the model, interpretability generally involves end-users in the process Mohseni et al. [2018], inducing strong biases. In fact, a qualitative evaluation alone cannot guarantee coherency to reality (i.e., model behavior), as false explanations can be more interpretable than accurate ones. To enable trust on XAI methods, we also need quantitative and objective evaluation metrics which validate the relation between the explanations produced by the XAI method and the trained model being assessed. The challenge of quantitatively evaluating XAI methods lies in the absence of a ground truth: we cannot be sure of what a DL method is doing unless we understand the model parametrization itself (at which point we would not need a XAI method). Nonetheless, we still want to validate the faithfulness Selvaraju et al. [2019] of XAI methods w.r.t. the underlying model, as this allows us to discern between accurate and misleading explanations.
arXiv.org Artificial Intelligence
Sep-28-2021