Whispers of Doubt Amidst Echoes of Triumph in NLP Robustness
Gupta, Ashim, Rajendhran, Rishanth, Stringham, Nathan, Srikumar, Vivek, Marasović, Ana
–arXiv.org Artificial Intelligence
Are the longstanding robustness issues in NLP resolved by today's larger and more performant models? To address this question, we conduct a thorough investigation using 19 models of different sizes spanning different architectural choices and pretraining objectives. We conduct evaluations using (a) OOD and challenge test sets, (b) CheckLists, (c) contrast sets, and (d) adversarial inputs. Our analysis reveals that not all OOD tests provide further insight into robustness. Evaluating with CheckLists and contrast sets shows significant gaps in model performance; merely scaling models does not make them sufficiently robust. Finally, we point out that current approaches for adversarial evaluations of models are themselves problematic: they can be easily thwarted, and in their current forms, do not represent a sufficiently deep probe of model robustness. We conclude that not only is the question of robustness in NLP as yet unresolved, but even some of the approaches to measure robustness need to be reassessed.
arXiv.org Artificial Intelligence
Nov-16-2023
- Country:
- Africa (0.93)
- Asia
- Middle East (1.00)
- Russia (0.68)
- Europe > United Kingdom (0.67)
- North America > United States
- California > San Francisco County
- San Francisco (0.14)
- Minnesota > Hennepin County
- Minneapolis (0.14)
- New Jersey > Essex County (0.14)
- Washington > King County
- Seattle (0.14)
- California > San Francisco County
- Genre:
- Personal (0.93)
- Research Report > New Finding (0.45)
- Industry:
- Leisure & Entertainment > Sports
- Baseball (0.92)
- Media (1.00)
- Banking & Finance (1.00)
- Government
- Foreign Policy (0.67)
- Military (1.00)
- Regional Government
- Asia Government (0.92)
- Europe Government (0.67)
- North America Government > United States Government (1.00)
- Transportation > Air (0.92)
- Law (1.00)
- Health & Medicine > Pharmaceuticals & Biotechnology (1.00)
- Information Technology > Security & Privacy (1.00)
- Law Enforcement & Public Safety > Crime Prevention & Enforcement (1.00)
- Leisure & Entertainment > Sports
- Technology: