Genre
Decision-Aligned Evaluation of Uncertainty Quantification
Schneider, Annika, Rochussen, Tommy, Stiller, Joshua, Fortuin, Vincent
Uncertainty estimates in machine learning are typically evaluated using generic metrics such as the negative log-likelihood and expected calibration error, yet good performance on such metrics does not necessarily imply high utility in downstream decisions. We introduce decision-alignment, a criterion that reveals which evaluation metrics meaningfully align with downstream utilities. Applying this framework, we show that many widely used uncertainty metrics are either misaligned with common decision problems or encode pathological prior beliefs about the downstream task. We then propose prior-weighted utility metrics, a special class of proper scoring rules that provides decision-aligned uncertainty evaluation. Across benchmark experiments and real-world case studies, our metrics consistently align with realized decision utility, while conventional metrics do not. Our results surface flaws in the current UQ evaluation protocol and offer a principled extension of existing metrics toward decision-relevant UQ evaluation.
$λ$-PSD: Scalable Approximate SNR-Optimised Polynomial Stein Discrepancies
Nguyen, Minh-Long, Vu, Thanh-Long, Drovandi, Christopher, South, Leah F., Nguyen, Trung-Tin
Polynomial Stein discrepancies (PSD) provide a scalable alternative to kernel Stein methods for measuring sample quality and goodness-of-fit testing, but their statistical properties remain poorly understood. We show that increasing polynomial degree primarily amplifies signal without adequately controlling variance, rather than directly optimising the signal-to-noise ratio (SNR). Under suitable assumptions, this might lead to a failure mode in which the $\text{SNR}^2$ can provably decay exponentially with polynomial degree. Motivated by this observation, we reformulate Stein discrepancy construction as an explicit $\text{SNR}^2$ maximisation problem, yielding a Rayleigh quotient over Stein features. This perspective motivates $λ$-PSD, an approximate scalable covariance-aware reweighting scheme defined in a low-dimensional subspace. Under Gaussian settings, we show that $λ$-PSD avoids the exponential $\text{SNR}^2$ collapse and achieves a stable $\text{SNR}^2$. Empirically, $λ$-PSD substantially improves test power while retaining linear-time complexity in the number of samples, highlighting the importance of SNR-aware design for scalable Stein discrepancies.
Ribbon: Scalable Approximation and Robust Uncertainty Quantification
Gibson, Graham, Tipton, John, Rumsey, Kellin, Klein, Natalie
Reliably quantifying predictive uncertainty is difficult for complex, high-dimensional, or misspecified models. Both fully Bayesian and bootstrap resampling methods provide principled uncertainty estimates but are often too expensive for modern machine-learning models because they require posterior sampling or repeated model refitting. We introduce Ribbon, a scalable approximation to Dirichlet-reweighted bootstrap uncertainty. Ribbon replaces repeated refitting with an influence-function linearization around a single fitted model, preserving the first-order data-reweighting structure of the Bayesian bootstrap while requiring only post-hoc linear algebra. Ribbon approximates the Bayesian-bootstrap or weighted-likelihood-bootstrap refitting target. With a general concentration parameter, Ribbon gives a calibrated Dirichlet-reweighting family whose uncertainty scale can be tuned on validation data. We show that Ribbon is asymptotically equivalent to a flat-prior Laplace approximation under correct likelihood specification and recovers the robust sandwich covariance under misspecification. Across synthetic regression, MNIST classification, and California Housing benchmarks, Ribbon provides competitive predictive performance and improved calibration in several settings while avoiding repeated model retraining.
Asymptotically Optimal Learning for Parametric Prophet Inequalities
Kim, Jung-hun, Grebennikova, Anna, Perchet, Vianney
We study learning in prophet inequalities with i.i.d. rewards drawn from an exponential-type parametric family with an unknown parameter $θ$, a class that includes exponential, Pareto, and bounded-support power-family distributions. We first characterize the optimal full-information asymptotic competitive ratio for this family. In the unbounded-support case, the limit is $ {\left(θ/({θ-c_+})\right)^{c_+/θ}}/ {Γ(1-c_+/θ)},$ while in the bounded-support case, the limit is $1$. We then propose a confidence-based dynamic-programming policy for online learning. By exploiting the explicit parametric structure, the policy achieves the same optimal asymptotic competitive ratio using only online observations, without external offline samples. We further derive distribution-specific convergence rates for canonical examples. Finally, numerical experiments on synthetic instances illustrate the performance of our algorithm.
When are likely answers right? On Sequence Probability and Correctness in LLMs
Zenn, Johannes, Geiping, Jonas
Many decoding methods for large language models can be understood as shifting probability mass toward outputs that are more likely under the model, either locally at the token level or globally at the sequence level. Therefore, their success depends on a fundamental question: when does sequence probability, that is, the conditional probability of a continuation given a prompt, actually align with correctness? In this paper, we set out to quantify this relationship across decoding methods, models, and benchmarks at four levels: across decoding methods, across hyperparameters within a method, across prompt-answer pairs within a dataset, and across repeated responses to the same prompt. We find that higher sequence probability is often predictive of correctness across prompt-answer pairs within a fixed dataset. However, this relationship does not generally transfer to decoding decisions: increasing sequence probability by changing hyperparameters or methods does not reliably improve accuracy. Further, sequence probability is not a good indicator of correctness for responses to the same prompt. These findings clarify when decoding can and cannot be expected to improve correctness, and provide practical guidance for decoding, self-consistency, and verifier-free self-improvement.
A probabilistic framework for online test-time adaptation
Corrales, Daniel, Insua, David Ríos
This paper presents a probabilistic framework for online test-time adaptation problems. In them, a model is trained on labeled data but must adapt to unlabeled data at test time under the assumption that training and test distributions potentially differ, that is, there might have been a distributional shift. The framework is based on a state-space modelling architecture from which parameter learning, parameter time evolution, prior tuning, and prediction can be characterized.
Robot dentist prepares tooth for a crown
The tiny robo-dentist will see you now. More information Adding us as a Preferred Source in Google by using this link indicates that you would like to see more of our content in Google News results. Its motors and control system are located outside the robot and connected to it via flexible drive shafts, cables, and tubes. Breakthroughs, discoveries, and DIY tips sent six days a week. By signing up, you confirm you are 16+, will receive newsletters and promotional content and agree to our Terms of Use and acknowledge the data practices in our Privacy Policy .
Hidden earthquake threat discovered beneath California could unleash devastating magnitude 7 tremor
Family secrets of Trump's closest White House aide: Natalie Harp's estranged socialist brother reveals ugly details of feud... and the tragedy that shattered everything Clay Aiken opens up about the'catastrophic' aftermath of grabbing Kelly Ripa's face live on air Iran's suicide drone strike on US ally threatens Trump's fragile peace in Strait of Hormuz Live, laugh, love mom charged with incestuous abuse of her two teenage adopted sons demands DIVORCE from handsome husband... as shocking custody request revealed Harry and Meghan may smugly believe the Establishment'plot' to return them is working. But they have no idea what Kate and William are thinking. My royal insiders have not held back... it's damning: RICHARD EDEN America's hottest housing market is a surprising East Coast city where 58% of homes sell above asking price Eva Longoria, 51, drops jaws in white string bikini as she displays gym-honed body during family beach day in Spain... after fleeing US Beloved Fox & Friends star says she's quitting after 22 years because of serious health condition Truth about Taylor Swift's'hookups' with ex Matty Healy... revealed by friends as his new model fiancée is accused of kinky pre-wedding stunts to humiliate Swift I lost a stone in 28 days WITHOUT weight-loss jabs: At size 32, I couldn't bear being fat anymore. This old-fashioned diet got me holiday-ready in weeks... YOU can do it too with these 6 steps Eerie'apocalyptic' sounds heard on Mount Shasta lead horseback riders to bizarre discovery Angelina Jolie war with Brad Pitt over sale of lavish French wine estate turns in Brad's favor as secretive vodka billionaire buyer forced to testify Lavish photos show Mark Zuckerberg's secretive new $170m hideout: First look behind the guarded gates of billionaire's palatial bunker Lionel Richie, 77, 'taken to hospital by ambulance' after dizzy spell onstage saw him end concert An anniversary present from Harry? Duchess of Sussex debuts new ring - and royal fans say it looks very similar to Kate's engagement band Grotesque'zombie squirrels' with oozing flesh pods spark alarm across the US BRYONY GORDON: Have you been tempted by the'Ozempic of alcohol' pill? I certainly was, but I've since faced a humbling truth.