On the Pitfalls of Heteroscedastic Uncertainty Estimation with Probabilistic Neural Networks

Seitzer, Maximilian, Tavakoli, Arash, Antic, Dimitrije, Martius, Georg

arXiv.org Machine Learning 

Capturing aleatoric uncertainty is a critical part of many machine learning systems. In deep learning, a common approach to this end is to train a neural network to estimate the parameters of a heteroscedastic Gaussian distribution by maximizing the logarithm of the likelihood function under the observed data. In this work, we examine this approach and identify potential hazards associated with the use of log-likelihood in conjunction with gradient-based optimizers. First, we present a synthetic example illustrating how this approach can lead to very poor but stable parameter estimates. Second, we identify the culprit to be the log-likelihood loss, along with certain conditions that exacerbate the issue. Third, we present an alternative formulation, termed β NLL, in which each data point's contribution to the loss is weighted by the β-exponentiated variance estimate. We show that using an appropriate β largely mitigates the issue in our illustrative example. Fourth, we evaluate this approach on a range of domains and tasks and show that it achieves considerable improvements and performs more robustly concerning hyperparameters, both in predictive RMSE and log-likelihood criteria. Endowing models with the ability to capture uncertainty is of crucial importance in machine learning. Uncertainty can be categorized into two main types: epistemic uncertainty and aleatoric uncertainty (Kiureghian & Ditlevsen, 2009). Epistemic uncertainty accounts for subjective uncertainty in the model, one that is reducible given sufficient data. By contrast, aleatoric uncertainty captures the stochasticity inherent in the observations and can itself be subdivided into homoscedastic and heteroscedastic uncertainty. Homoscedastic uncertainty corresponds to noise that is constant across the input space, whereas heteroscedastic uncertainty corresponds to noise that varies with the input. There are well-established benefits for modeling each type of uncertainty.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found