Faithfulness and the Notion of Adversarial Sensitivity in NLP Explanations

Oct-9-2024–arXiv.org Artificial Intelligence

Faithfulness is arguably the most critical metric to assess the reliability of explainable AI. In NLP, current methods for faithfulness evaluation are fraught with discrepancies and biases, often failing to capture the true reasoning of models. We introduce Adversarial Sensitivity as a novel approach to faithfulness evaluation, focusing on the explainer's response when the model is under adversarial attack. Our method accounts for the faithfulness of explainers by capturing sensitivity to adversarial input changes. This work addresses significant limitations in existing evaluation techniques, and furthermore, quantifies faithfulness from a crucial yet underexplored paradigm.

explainer, machine learning, natural language, (16 more...)

arXiv.org Artificial Intelligence

Oct-9-2024

arXiv.org PDF

Add feedback

Genre:
- Research Report > New Finding (0.46)

Industry:
- Information Technology > Security & Privacy (0.48)
- Government > Military (0.34)

Technology:
- Information Technology > Artificial Intelligence
  - Natural Language > Explanation & Argumentation (0.48)
  - Representation & Reasoning > Search (0.46)
  - Machine Learning
    - Statistical Learning (0.68)
    - Neural Networks (0.47)

Duplicate Docs Excel Report

Title
None found

Similar Docs Excel Report more

Title	Similarity	Source
None found