Beyond Interpretability: The Gains of Feature Monosemanticity on Model Robustness

Open in new window