Domain Classification-based Source-specific Term Penalization for Domain Adaptation in Hate-speech Detection
Bose, Tulika, Aletras, Nikolaos, Illina, Irina, Fohr, Dominique
–arXiv.org Artificial Intelligence
State-of-the-art approaches for hate-speech detection usually exhibit poor performance in out-of-domain settings. This occurs, typically, due to classifiers overemphasizing source-specific information that negatively impacts its domain invariance. Prior work has attempted to penalize terms related to hate-speech from manually curated lists using feature attribution methods, which quantify the importance assigned to input terms by the classifier when making a prediction. We, instead, propose a domain adaptation approach that automatically extracts and penalizes source-specific terms using a domain classifier, which learns to differentiate between domains, and feature-attribution scores for hate-speech classes, yielding consistent improvements in cross-domain evaluation.
arXiv.org Artificial Intelligence
Sep-18-2022
- Country:
- Asia > China (0.04)
- Oceania > Australia
- North America
- Mexico (0.04)
- Canada (0.04)
- United States
- New York > New York County
- New York City (0.04)
- Minnesota > Hennepin County
- Minneapolis (0.14)
- Florida > Palm Beach County
- Boca Raton (0.04)
- California > San Diego County
- San Diego (0.04)
- New York > New York County
- Europe
- Spain (0.04)
- United Kingdom > England
- South Yorkshire > Sheffield (0.04)
- Italy
- Tuscany > Florence (0.04)
- Marche > Ancona Province
- Ancona (0.04)
- Ireland > Leinster
- County Dublin > Dublin (0.04)
- France
- Belgium > Brussels-Capital Region
- Brussels (0.04)
- Genre:
- Research Report (1.00)
- Industry:
- Government (0.47)
- Leisure & Entertainment (0.46)
- Law Enforcement & Public Safety > Crime Prevention & Enforcement (0.46)
- Technology: