Understanding and Mitigating Spurious Correlations in Text Classification with Neighborhood Analysis

Chew, Oscar, Lin, Hsuan-Tien, Chang, Kai-Wei, Huang, Kuan-Hao

Jul-16-2023–arXiv.org Artificial Intelligence

Recent research has revealed that deep learning models have a tendency to leverage spurious correlations that exist in the training set but may not hold true in general circumstances. For instance, a sentiment classifier may erroneously learn that the token performances is commonly associated with positive movie reviews. Relying on these spurious correlations degrades the classifiers performance when it deploys on out-of-distribution data. In this paper, we examine the implications of spurious correlations through a novel perspective called neighborhood analysis. The analysis uncovers how spurious correlations lead unrelated words to erroneously cluster together in the embedding space. Driven by the analysis, we design a metric to detect spurious tokens and also propose a family of regularization methods, NFL (doN't Forget your Language) to mitigate spurious correlations in text classification. Experiments show that NFL can effectively prevent erroneous clusters and significantly improve the robustness of classifiers.

correlation, machine learning, natural language, (20 more...)

arXiv.org Artificial Intelligence

Jul-16-2023

arXiv.org PDF

Add feedback

Country:
- South America > Brazil (0.04)
- Oceania > Australia
  - Victoria > Melbourne (0.04)
- North America
  - Canada (0.04)
  - Dominican Republic (0.04)
  - United States
    - Washington > King County
      - Seattle (0.14)
    - Minnesota > Hennepin County
      - Minneapolis (0.14)
    - Louisiana > Orleans Parish
      - New Orleans (0.04)
    - California > Los Angeles County
      - Los Angeles (0.14)
- Europe
  - United Kingdom (0.04)
  - Spain (0.04)
  - Germany (0.04)
  - France (0.04)
  - Italy > Tuscany
    - Florence (0.04)
  - Ireland > Leinster
    - County Dublin > Dublin (0.04)
  - Belgium > Brussels-Capital Region
    - Brussels (0.04)
- Asia
  - China > Hong Kong (0.04)
  - Taiwan (0.04)
  - Middle East
    - Republic of Türkiye (0.04)
    - UAE > Abu Dhabi Emirate
      - Abu Dhabi (0.04)

Genre:
- Research Report (0.82)

Industry:
- Media > Film (0.34)
- Leisure & Entertainment (0.34)

Technology:
- Information Technology > Artificial Intelligence
  - Natural Language > Text Classification (0.71)
  - Machine Learning > Neural Networks
    - Deep Learning (0.68)

Duplicate Docs Excel Report

Title
None found

Similar Docs Excel Report more

Title	Similarity	Source
None found