Detecting and Understanding Vulnerabilities in Language Models via Mechanistic Interpretability

Open in new window