Decomposing MLP Activations into Interpretable Features via Semi-Nonnegative Matrix Factorization

Open in new window