Disentangling Recall and Reasoning in Transformer Models through Layer-wise Attention and Activation Analysis

Open in new window