Representation Disentaglement via Regularization by Identification
–arXiv.org Artificial Intelligence
This work focuses on the problem of learning disentangled representations from observational data. Given observations ${\mathbf{x}^{(i)}}$ for $i=1,...,N $ drawn from $p(\mathbf{x}|\mathbf{y})$ with generative variables $\mathbf{y}$ admitting the distribution factorization $p(\mathbf{y}) = \prod_{c} p(\mathbf{y}_c )$, we ask whether learning disentangled representations matching the space of observations with identification guarantees on the posterior $p(\mathbf{z}| \mathbf{x}, \hat{\mathbf{y}}_c)$ for each $c$, is plausible. We argue modern deep representation learning models of data matching the distributed factorization property are ill-posed with collider bias behaviour; a source of bias producing entanglement between generating variables. Under the rubric of causality, we show this issue can be explained and reconciled under the condition of identifiability; attainable under supervision or a weak-form of it. For this, we propose regularization by identification (ReI), a modular regularization engine designed to align the behavior of large scale DL models with domain knowledge. Empirical evidence shows that enforcing ReI in a variational framework results in interpretable disentangled representations equipped with generalization capabilities to out-of-distribution examples and that aligns nicely with the true expected effect from domain knowledge between generating variables and measurement apparatus.
arXiv.org Artificial Intelligence
Jun-15-2023
- Country:
- North America > United States
- Virginia > Arlington County
- Arlington (0.04)
- New Mexico > Los Alamos County
- Los Alamos (0.04)
- Virginia > Arlington County
- Europe > United Kingdom
- England > Cambridgeshire > Cambridge (0.04)
- North America > United States
- Genre:
- Research Report (0.82)
- Industry:
- Health & Medicine (1.00)
- Technology: