Uncertainty
Gaussian Mean Field Variational Inference can Overestimate Predictive Variance
Odgers, James, Riegler, Ben, Swaroop, Siddharth, Fortuin, Vincent
Mean Field Variational Inference (MFVI) is widely understood to underestimate posterior variance. By analysing conjugate Bayesian Linear Regression (BLR), we show that this characterization is incomplete: while MFVI underestimates the variance in parameter space, it can overestimate the predictive variance compared to the exact posterior. We show that if the MFVI posterior underestimates predictive variances in some directions, it necessarily overestimates them in others. Crucially, this overestimation occurs in directions where the training data concentrates. This leads to the surprising result that, for a test point drawn from the training distribution, MFVI's expected predictive variance exceeds that of the exact posterior. We demonstrate a pathological case of this effect, where the MFVI posterior fails to reduce predictive variance compared to the prior on in distribution data. We connect these results to the Cold Posterior Effect, arguing that varying the temperature can correct this overestimation, yielding predictions closer to those of the exact posterior. We validate our theory on synthetic and real-world regression tasks.
Stochastic Expectation Maximization for Robust State-Space Radio Interferometric Imaging
Arab, Nawel, Korso, Mohammed Nabil El, Vin, Isabelle, Larzabal, Pascal
State-space models provide a powerful framework for describing the evolution of hidden states in dynamical systems [3], [4], [1]. Conventionally, state-space models assume Gaussian measurement and state noise, owing to their tractability and well-characterized statistical properties. However, many real-world phenomena are subject to perturbations that deviate from the conventional Gaussian noise assumption. In radio interferometry, for instance, observational data are frequently corrupted by non-Gaussian noise sources such as radio-frequency interference (RFI) [5], [2], which originates from man-made signals and introduces significant distortions into astronomical measurements [6], [30]. Such interference produces sporadic high-power spikes in the measured visibilities, leading to heavy-tailed statistics. Many radio-interferometric reconstruction methods assume Gaussian additive noise [7], [31], [33], [35], an approximation that can lead to inaccurate reconstructions when the heavy-tailed nature of real-world measurement noise is not properly accounted for. In the realm of state-space modeling, addressing non-Gaussian noise has led to the development of various methodological approaches, notably particle filtering and non-conventional Kalman filters. Particle filters [8], or Sequential Monte Carlo methods, are designed to handle non-linear and non-Gaussian state-space models by representing the posterior distribution with a set of weighted samples [9], [10], [32].
Tensor-based second-order causal discovery
Ouyang, Nathan, Wang, Kexin, Seigal, Anna
Causal discovery seeks to uncover the causal dependencies among variables. For this purpose, we propose an algorithm called Tensor-based Second-order Causal Discovery (TSCD). Its input is a tensor obtained from the covariance matrices of observational and interventional data. Assuming the causal dependencies follow a linear structural equation model on a directed acyclic graph (DAG), TSCD outputs the DAG and the functions on its edges, requiring only that the noise variables are uncorrelated. We also implement a version of the approach for nonlinear models. Our focus on second-order statistics (via the covariance matrices) is motivated by their statistical and computational efficiency relative to higher-order moments, their identifiability relative to first-order statistics, and that they work regardless of whether the variables are Gaussian. We show that TSCD has identifiable causal order and parameters from a number of interventions that is logarithmic in the number of variables. Experiments show that TSCD is robust to noise, competitive with existing methods, and scales to hundreds of variables.
A Step Towards Inherently Interpretable Causal Machine Learning Models For Decision Support
The growing reliance on machine learning for decisions across sectors underscores the importance of model transparency and interpretability. Existing post-hoc explainability methods and inherently interpretable approaches shed light on model behavior, yet they primarily reveal how models exploit correlations to maximize performance in prediction tasks. However, many decisions require causal insights and the possibility of using models for what-if scenario evaluation. To address this, we propose the integration of causal machine learning with inherently interpretable models for cross-sectional data. We evaluate these methods in terms of predictive accuracy and interpretability. Our findings show that the proposed approach achieves competitive performance in prediction and what-if analysis while offering transparency on the system structure, causal relationships among variables, and the functional forms that connect them. This work contributes to research on causality, machine learning interpretability, and data-driven decision support by offering informed, transparent, and causally grounded decisions.
Meta-D2AG: Causal Graph Learning with Interventional Dynamic Data
Causal discovery in the form of a directed acyclic graph (DAG) for dynamic time series data has been widely studied in various applications. In this work, we propose a dynamic DAG discovery algorithm, Meta-D2AG, based on online metalearning. Meta-D2AG is designed to learn dynamic DAG structures from potentially nonlinear and non-stationary time series datasets, accounting for changes in both parameters and graph structures. Unlike most of the existing work focusing on observational, offline, and/or stationary settings, Meta-D2AG explicitly treats data collected at different time points with distribution shifts as distinct domains, which is assumed to occur as a result of external interventions. Moreover, MetaD2AG involves a new online meta-learning framework to take advantage of the temporal transition among existing domains such that it can quickly adapt to new domains with few measurements. A first-order optimization approach is utilized to efficiently solve the meta-learning framework, and theoretical analysis establishes the identifiability conditions and the convergence of the learning process. We demonstrate the promising performance of the proposed meta learning framework through better accuracy on benchmark datasets against state-of-the-art baselines.
Wonder Wins Ways: Curiosity-Driven Exploration through Multi-Agent Contextual Calibration
Autonomous exploration in complex multi-agent reinforcement learning (MARL) with sparse rewards critically depends on providing agents with effective intrinsic motivation. While artificial curiosity offers a powerful self-supervised signal, it often confuses environmental stochasticity with meaningful novelty.
On the Global Optimality of Policy Gradient Methods in General Utility Reinforcement Learning
Reinforcement learning with general utilities (RLGU) offers a unifying framework to capture several problems beyond standard expected returns, including imitation learning, pure exploration, and safe RL. Despite recent fundamental advances in the theoretical analysis of policy gradient (PG) methods for standard RL and recent efforts in RLGU, the understanding of these PG algorithms and their scope of application in RLGU still remain limited. In this work, we establish global optimality guarantees of PG methods for RLGU in which the objective is a general concave utility function of the state-action occupancy measure. In the tabular setting, we provide global optimality results using a new proof technique building on recent theoretical developments on the convergence of PG methods for standard RL using gradient domination. Our proof technique opens avenues for analyzing policy parameterizations beyond the direct policy parameterization for RLGU. In addition, we provide global optimality results for large state-action space settings beyond prior work which has mostly focused on the tabular setting. In this large scale setting, we adapt PG methods by approximating occupancy measures within a function approximation class using maximum likelihood estimation. Our sample complexity only scales with the dimension induced by our approximation class instead of the size of the state-action space.