Causal Intervention Framework for Variational Auto Encoder Mechanistic Interpretability

Roy, Dip

arXiv.org Artificial Intelligence 

To establish a basis for understanding how generative models represent and transform data is an essential problem in the field of deep learning interpretation. Although, the mechanistic interpretation of discriminative architectures (e.g., transformers) has produced substantial new insights about these types of systems, there has been relatively little work on mechanistic interpretation of variational autoencoders (VAEs). VAEs are widely used for representation learning; this paper presents the first general-purpose multilevel causal intervention framework for the mechanistic interpretation of VAEs. The framework includes four basic manipulation types: input manipulation, latent-space perturbation, activation-patching, and causal-mediation-analysis. This paper also defines three new quantitative metrics that measure various characteristics of VAE internal representations that are not measured by existing disentanglement metrics alone: CausalEffect-Strength (CES), intervention specificity, and circuit modularity. We conducted the largest empirical study to date of the causal mechanisms of six VAE architectures (standard VAE, β-VAE, FactorVAE, β-TC-VAE, DIP-VAE-II, and VQ-VAE) over five different benchmark datasets (dSprites, 3DShapes, MPI3D, CelebA, and SmallNORB) using three different random seeds per architecture and dataset pair, leading to a total of 90 independent training runs. Our empirical results reveal several new findings: (i) a consistent within-dataset negative correlation between CES and DCI disentanglement, which we refer to as the CES-DCI trade-off; (ii) that the KL reweighting mechanism of β-VAE can cause a capacity bottleneck for β-VAE when the number of generative factors in a model approaches its latent dimensionality, causing a degradation of disentanglement performance on complex datasets; (iii) that no single VAE architecture will be best for all of the five datasets tested -- the optimal choice of VAE architecture is dependent upon the structure of the dataset being modeled; and (iv) that CES-based metrics applied to discrete latent spaces (VQ-VAE) yield near zero values, indicating a critical limitation of continuousintervention methods for analyzing discrete latent spaces. These results provide both a theoretical foundation and empirical evaluations of the interpretation of generative models.