The Loss Kernel: A Geometric Probe for Deep Learning Interpretability
Adam, Maxwell, Furman, Zach, Hoogland, Jesse
–arXiv.org Artificial Intelligence
A central goal in AI interpretability and data attribution is interpreting and mapping the global structure of the data distribution as seen by a trained neural network (Carter et al., 2019; Pepin Lehalleur et al., 2025; Olah, 2015). One approach is to start local, by quantifying a suitable measure of similarity between pairs of individual samples--that is, by defining a kernel. "Interpreting" the global structure of the data distribution then becomes a problem of analyzing the geometric structure in this kernel (e.g., via clustering techniques), and "mapping" becomes a problem of visualizing points in this kernel space (e.g., via dimensionality reduction techniques). This kernel-based approach has been used successfully with similarity measures derived from activations or representations. For example, it is possible to define a kernel via cosine similarity between the hidden vectors of sparse autoencoders (SAEs).
arXiv.org Artificial Intelligence
Oct-1-2025