Country
meval: A Statistical Toolbox for Fine-Grained Model Performance Analysis
Sutariya, Dishantkumar, Petersen, Eike
Analyzing machine learning model performance stratified by patient and recording properties is becoming the accepted norm and often yields crucial insights about important model failure modes. Performing such analyses in a statistically rigorous manner is non-trivial, however. Appropriate performance metrics must be selected that allow for valid comparisons between groups of different sample sizes and base rates; metric uncertainty must be determined and multiple comparisons be corrected for, in order to assess whether any observed differences may be purely due to chance; and in the case of intersectional analyses, mechanisms must be implemented to find the most `interesting' subgroups within combinatorially many subgroup combinations. We here present a statistical toolbox that addresses these challenges and enables practitioners to easily yet rigorously assess their models for potential subgroup performance disparities. While broadly applicable, the toolbox is specifically designed for medical imaging applications. The analyses provided by the toolbox are illustrated in two case studies, one in skin lesion malignancy classification on the ISIC2020 dataset and one in chest X-ray-based disease classification on the MIMIC-CXR dataset.
Sharp Structure-Agnostic Lower Bounds for General Functional Estimation
Jin, Jikai, Syrgkanis, Vasilis
The design of efficient nonparametric estimators has long been a central problem in statistics, machine learning, and decision making. Classical optimal procedures often rely on strong structural assumptions, which can be misspecified in practice and complicate deployment. This limitation has sparked growing interest in structure-agnostic approaches -- methods that debias black-box nuisance estimates without imposing structural priors. Understanding the fundamental limits of these methods is therefore crucial. This paper provides a systematic investigation of the optimal error rates achievable by structure-agnostic estimators. We first show that, for estimating the average treatment effect (ATE), a central parameter in causal inference, doubly robust learning attains optimal structure-agnostic error rates. We then extend our analysis to a general class of functionals that depend on unknown nuisance functions and establish the structure-agnostic optimality of debiased/double machine learning (DML). We distinguish two regimes -- one where double robustness is attainable and one where it is not -- leading to different optimal rates for first-order debiasing, and show that DML is optimal in both regimes. Finally, we instantiate our general lower bounds by deriving explicit optimal rates that recover existing results and extend to additional estimands of interest. Our results provide theoretical validation for widely used first-order debiasing methods and guidance for practitioners seeking optimal approaches in the absence of structural assumptions. This paper generalizes and subsumes the ATE lower bound established in \citet{jin2024structure} by the same authors.
Penalized Fair Regression for Multiple Groups in Chronic Kidney Disease
Nakamoto, Carter H., Chen, Lucia Lushi, Foryciarz, Agata, Rose, Sherri
Fair regression methods have the potential to mitigate societal bias concerns in health care, but there has been little work on penalized fair regression when multiple groups experience such bias. We propose a general regression framework that addresses this gap with unfairness penalties for multiple groups. Our approach is demonstrated for binary outcomes with true positive rate disparity penalties. It can be efficiently implemented through reduction to a cost-sensitive classification problem. We additionally introduce novel score functions for automatically selecting penalty weights. Our penalized fair regression methods are empirically studied in simulations, where they achieve a fairness-accuracy frontier beyond that of existing comparison methods. Finally, we apply these methods to a national multi-site primary care study of chronic kidney disease to develop a fair classifier for end-stage renal disease. There we find substantial improvements in fairness for multiple race and ethnicity groups who experience societal bias in the health care system without any appreciable loss in overall fit.
Smoothing DiLoCo with Primal Averaging for Faster Training of LLMs
Defazio, Aaron, Mishchenko, Konstantin, Raman, Parameswaran, Shi, Hao-Jun Michael, Xiao, Lin
We propose Generalized Primal Averaging (GPA), an extension of Nesterov's method in its primal averaging formulation that addresses key limitations of recent averaging-based optimizers such as single-worker DiLoCo and Schedule-Free (SF) in the non-distributed setting. These two recent algorithmic approaches improve the performance of base optimizers, such as AdamW, through different iterate averaging strategies. Schedule-Free explicitly maintains a uniform average of past weights, while single-worker DiLoCo performs implicit averaging by periodically aggregating trajectories, called pseudo-gradients, to update the model parameters. However, single-worker DiLoCo's periodic averaging introduces a two-loop structure, increasing its memory requirements and number of hyperparameters. GPA overcomes these limitations by decoupling the interpolation constant in the primal averaging formulation of Nesterov. This decoupling enables GPA to smoothly average iterates at every step, generalizing and improving upon single-worker DiLoCo. Empirically, GPA consistently outperforms single-worker DiLoCo while removing the two-loop structure, simplifying hyperparameter tuning, and reducing its memory overhead to a single additional buffer. On the Llama-160M model, GPA provides a 24.22% speedup in terms of steps to reach the baseline (AdamW's) validation loss. Likewise, GPA achieves speedups of 12% and 27% on small and large batch setups, respectively, to attain AdamW's validation accuracy on the ImageNet ViT workload. Furthermore, we prove that for any base optimizer with regret bounded by $O(\sqrt{T})$, where $T$ is the number of iterations, GPA can match or exceed the convergence guarantee of the original optimizer, depending on the choice of interpolation constants.
Safeguarded Stochastic Polyak Step Sizes for Non-smooth Optimization: Robust Performance Without Small (Sub)Gradients
Oikonomou, Dimitris, Loizou, Nicolas
The stochastic Polyak step size (SPS) has proven to be a promising choice for stochastic gradient descent (SGD), delivering competitive performance relative to state-of-the-art methods on smooth convex and non-convex optimization problems, including deep neural network training. However, extensions of this approach to non-smooth settings remain in their early stages, often relying on interpolation assumptions or requiring knowledge of the optimal solution. In this work, we propose a novel SPS variant, Safeguarded SPS (SPS$_{safe}$), for the stochastic subgradient method, and provide rigorous convergence guarantees for non-smooth convex optimization with no need for strong assumptions. We further incorporate momentum into the update rule, yielding equally tight theoretical results. On non-smooth convex benchmarks, our experiments are consistent with the theoretical predictions on how the safeguard affects the convergence neighborhood. On deep neural networks the proposed step size achieves competitive performance to existing adaptive baselines and exhibits stable behavior across a wide range of problem settings. Moreover, in these experiments, the gradient norms under our step size do not collapse to (near) zero, indicating robustness to vanishing gradients.
Mass power outages affect 130,000 in San Francisco and disrupt traffic
A widespread power failure plunged San Francisco into darkness on Saturday night, disrupting traffic citywide and forcing numerous self-driving Waymo taxis to stop abruptly in the middle of streets and intersections. As electricity went out across large portions of the city, traffic signals failed, leaving autonomous vehicles unable to operate as normal. Photos and videos shared by users on X showed Waymo robotaxis frozen in place, backing up traffic and creating hazardous conditions for other drivers. Waymo confirmed on Saturday evening that it had shut down its driverless ride-hailing service throughout San Francisco after footage circulated online showing its vehicles blocking roads during the blackout. "We have temporarily suspended our ride-hailing services in the San Francisco Bay Area due to the widespread power outage," Waymo spokesperson Suzanne Philion said in a statement to several news outlets.
More than 20,000 still without power after massive San Francisco blackout
Things to Do in L.A. Tap to enable a layout that focuses on the article. This is read by an automated voice. Please report any issues or inconsistencies here . After Saturday's blackout, roughly 110,000 San Francisco residents have power again. About 21,000 are still in the dark as extensive repairs continue after a substation fire.
A San Francisco power outage left Waymo's self-driving cars stranded at intersections
LG TVs add'delete' option for Copilot A San Francisco power outage left Waymo's self-driving cars stranded at intersections Waymo halted its autonomous ride-hailing services in the city in response. Several of Waymo's autonomous vehicles were seen stuck in the middle of San Francisco streets following a significant power outage that took out the city's traffic lights. Waymo responded to the power outage by suspending its ride-hailing services in the city, but images and videos on social media showed the self-driving taxis stopped at intersections with hazard lights on. We have temporarily suspended our ride-hailing services in the San Francisco Bay Area due to the widespread power outage, Suzanne Philion, a spokesperson for Waymo, told Engadget in an email. Our teams are working diligently and in close coordination with city officials, and we are hopeful to bring our services back online soon.
9 new butterflies discovered in old museum archives
The team even extracted DNA from a tiny 100-year-old butterfly leg. Breakthroughs, discoveries, and DIY tips sent every weekday. When you think of butterflies, chances are you imagine unmistakable insects with bright, bold wings. But it turns out that individual butterfly species are sometimes shockingly difficult to tell apart. "Thanks to the genetic revolution and the collaboration of researchers and museums in various countries led by London's Natural History Museum, century-old butterflies are now speaking to us," Christophe Faynel, an entomologist at the Société entomologique Antilles Guyane, said in a statement .
Alcohol consumption falls to a record low in Britain - so, do you drink more or less than the national average?
SNL savages Trump after releasing the Epstein files in cold open... but MAGA have the last laugh Retirees are ditching golf and sun for this unlikely city...as top destinations revealed Frail woman found bludgeoned to death next to'bloodied skateboard' in her NYC apartment You've only been told half the story about the Reiner murders. These hidden horrors MUST be outed... before Hollywood's sick secrecy pact wins: MAUREEN CALLAHAN Monumental downfall of it-girl streamer who snubbed Drake's advances... as she's engulfed by disgusting scandals and family tragedy Devastating truth about Rob Reiner's daughter Romy: Her own addiction battle... how she'lived in fear' of Nick... and the handsome companion she's leaning on, all revealed by heartbroken friends Smellebrities! 10 actors whose hygiene habits prompted complaints from co-stars - after Charlotte Church said she's stopped wearing deodorant and'generally stinks' Crime drama dubbed'the greatest of all time' with the'perfect ending' and a whopping 96% Rotten Tomatoes score is now free to stream - plus there's a reboot in the works Then I learned what he told her about our sex life and I can't even look at him: ASK JANA America's cutest Christmas village battles to save itself after being hit by storms and severe flooding Prince Harry's controversial comment at Christmas that sparked Meghan Markle's bitter family feud Father Christmas is'too white' and has no right to judge if children are naughty or nice, says woke museum Inside Tinseltown's'cursed' neighborhood where Rob Reiner was murdered that also saw Marilyn Monroe and Nicole Brown Simpson's deaths Deputy Attorney General reveals REAL reason why Trump's picture in Epstein files was taken down: 'It's absurd and laughable' How Tom Brady REALLY feels about Gisele Bundchen's secret wedding to jiu-jitsu instructor... as insiders whisper about potential of his OWN second marriage Alcohol consumption falls to a record low in Britain - so, do you drink more or less than the national average? READ MORE: Do you drink more than your partner? It's a typically boozy time of year - but Brits are drinking less alcohol than in decades gone by, according to new figures. Data released by research company IWSR reveals the average UK adult consumed 10.2 alcoholic drinks a week last year.