A Comparative Analysis of Sparse Autoencoder and Activation Difference in Language Model Steering

Open in new window