Toward a Theory of Generalizability in LLM Mechanistic Interpretability Research

Open in new window