Goto

Collaborating Authors

 Industry


Communication-Efficient Federated Risk Difference Estimation for Time-to-Event Clinical Outcomes

arXiv.org Machine Learning

Privacy-preserving model co-training in medical research is often hindered by server-dependent architectures incompatible with protected hospital data systems and by the predominant focus on relative effect measures (hazard ratios) which lack clinical interpretability for absolute survival risk assessment. We propose FedRD, a communication-efficient framework for federated risk difference estimation in distributed survival data. Unlike typical federated learning frameworks (e.g., FedAvg) that require persistent server connections and extensive iterative communication, FedRD is server-independent with minimal communication: one round of summary statistics exchange for the stratified model and three rounds for the unstratified model. Crucially, FedRD provides valid confidence intervals and hypothesis testing--capabilities absent in FedAvg-based frameworks. We provide theoretical guarantees by establishing the asymptotic properties of FedRD and prove that FedRD (unstratified) is asymptotically equivalent to pooled individual-level analysis. Simulation studies and real-world clinical applications across different countries demonstrate that FedRD outperforms local and federated baselines in both estimation accuracy and prediction performance, providing an architecturally feasible solution for absolute risk assessment in privacy-restricted, multi-site clinical studies.


Optimality of Staircase Mechanisms for Vector Queries under Differential Privacy

arXiv.org Machine Learning

We study the optimal design of additive mechanisms for vector-valued queries under $ฮต$-differential privacy (DP). Given only the sensitivity of a query and a norm-monotone cost function measuring utility loss, we ask which noise distribution minimizes expected cost among all additive $ฮต$-DP mechanisms. Using convex rearrangement theory, we show that this infinite-dimensional optimization problem admits a reduction to a one-dimensional compact and convex family of radially symmetric distributions whose extreme points are the staircase distributions. As a consequence, we prove that for any dimension, any norm, and any norm-monotone cost function, there exists an $ฮต$-DP staircase mechanism that is optimal among all additive mechanisms. This result resolves a conjecture of Geng, Kairouz, Oh, and Viswanath, and provides a geometric explanation for the emergence of staircase mechanisms as extremal solutions in differential privacy.


engGNN: A Dual-Graph Neural Network for Omics-Based Disease Classification and Feature Selection

arXiv.org Machine Learning

Omics data, such as transcriptomics, proteomics, and metabolomics, provide critical insights into disease mechanisms and clinical outcomes. However, their high dimensionality, small sample sizes, and intricate biological networks pose major challenges for reliable prediction and meaningful interpretation. Graph Neural Networks (GNNs) offer a promising way to integrate prior knowledge by encoding feature relationships as graphs. Yet, existing methods typically rely solely on either an externally curated feature graph or a data-driven generated one, which limits their ability to capture complementary information. To address this, we propose the external and generated Graph Neural Network (engGNN), a dual-graph framework that jointly leverages both external known biological networks and data-driven generated graphs. Specifically, engGNN constructs a biologically informed undirected feature graph from established network databases and complements it with a directed feature graph derived from tree-ensemble models. This dual-graph design produces more comprehensive embeddings, thereby improving predictive performance and interpretability. Through extensive simulations and real-world applications to gene expression data, engGNN consistently outperforms state-of-the-art baselines. Beyond classification, engGNN provides interpretable feature importance scores that facilitate biologically meaningful discoveries, such as pathway enrichment analysis. Taken together, these results highlight engGNN as a robust, flexible, and interpretable framework for disease classification and biomarker discovery in high-dimensional omics contexts.


Large Data Limits of Laplace Learning for Gaussian Measure Data in Infinite Dimensions

arXiv.org Machine Learning

Laplace learning is a semi-supervised method, a solution for finding missing labels from a partially labeled dataset utilizing the geometry given by the unlabeled data points. The method minimizes a Dirichlet energy defined on a (discrete) graph constructed from the full dataset. In finite dimensions the asymptotics in the large (unlabeled) data limit are well understood with convergence from the graph setting to a continuum Sobolev semi-norm weighted by the Lebesgue density of the data-generating measure. The lack of the Lebesgue measure on infinite-dimensional spaces requires rethinking the analysis if the data aren't finite-dimensional. In this paper we make a first step in this direction by analyzing the setting when the data are generated by a Gaussian measure on a Hilbert space and proving pointwise convergence of the graph Dirichlet energy.


When Are Two Scores Better Than One? Investigating Ensembles of Diffusion Models

arXiv.org Machine Learning

Diffusion models now generate high-quality, diverse samples, with an increasing focus on more powerful models. Although ensembling is a well-known way to improve supervised models, its application to unconditional score-based diffusion models remains largely unexplored. In this work we investigate whether it provides tangible benefits for generative modelling. We find that while ensembling the scores generally improves the score-matching loss and model likelihood, it fails to consistently enhance perceptual quality metrics such as FID on image datasets. We confirm this observation across a breadth of aggregation rules using Deep Ensembles, Monte Carlo Dropout, on CIF AR-10 and FFHQ. We attempt to explain this discrepancy by investigating possible explanations, such as the link between score estimation and image quality. We also look into tabular data through random forests, and find that one aggregation strategy outperforms the others. Finally, we provide theoretical insights into the summing of score models, which shed light not only on ensembling but also on several model composition techniques (e.g.


Physics-Informed Singular-Value Learning for Cross-Covariances Forecasting in Financial Markets

arXiv.org Machine Learning

A new wave of work on covariance cleaning and nonlinear shrinkage has delivered asymptotically optimal analytical solutions for large covariance matrices. The same framework has been generalized to empirical cross-covariance matrices, whose singular value decomposition identifies canonical comovement modes between two asset sets, with singular values quantifying the strength of each mode and providing natural targets for shrinkage. Existing analytical cross-covariance cleaners are derived under strong stationarity and large-sample assumptions, and they typically rely on mesoscopic regularity conditions such as bounded spectra; macroscopic common modes (e.g., a global market factor) violate these conditions. When applied to real equity returns, where dependence structures drift over time and global modes are prominent, we find that these theoretically optimal formulas do not translate into robust out-of-sample performance. We address this gap by designing a random-matrix-inspired neural architecture that operates in the empirical singular-vector basis and learns a nonlinear mapping from empirical singular values to their corresponding cleaned values. By construction, the network can recover the analytical solution as a special case, yet it remains flexible enough to adapt to non-stationary dynamics and mode-driven distortions. Trained on a long history of equity returns, the proposed method achieves a more favorable bias-variance trade-off than purely analytical cleaners and delivers systematically lower out-of-sample cross-covariance prediction errors. Our results demonstrate that combining random-matrix theory with machine learning makes asymptotic theories practically effective in realistic time-varying markets.


US House panel advances bill to give Congress authority on AI chip exports

Al Jazeera

What is the Insurrection Act? Why is the US Fed chair criminal probe causing alarm? The United States House of Representatives Foreign Affairs Committee has overwhelmingly voted to advance a bill that would give Congress more power over artificial intelligence chip exports despite pushback from White House AI tsar David Sacks and a social media campaign against the legislation. Representative Brian Mast of Florida, a Republican and the chair of the House Foreign Affairs Committee, introduced the "AI Overwatch Act" in December after US President Donald Trump greenlit shipments of Nvidia's powerful H200 AI chips to China. The bill claims that those "countries of concern" also include countries beyond China, such as Russia, Iran, North Korea, Cuba and Venezuela.


Apple is reportedly overhauling Siri to be an AI chatbot

Engadget

Bungie's Marathon arrives on March 5 How to claim Verizon's $20 outage credit This new Gemini-powered approach to Siri could be coming in 2027. Apple has been spinning its wheels for many months over its approach to artificial intelligence, but a strategy finally appears to be emerging for the company. 's Mark Gurman reported today that Apple's long-awaited Siri overhaul will allegedly involve transforming the voice assistant into an AI chatbot, internally called Campos. Sources have reportedly told Gurman that Apple chatbot will completely replace the current Siri interface in favor of a more interactive model similar to those used by OpenAI's ChatGPT and Google's Gemini. He also cited sources who claimed that while Apple has been testing a standalone Campos app, the company doesn't plan to release it for customers.


Apple is reportedly developing a wearable AI pin

Engadget

Bungie's Marathon arrives on March 5 How to claim Verizon's $20 outage credit The device is said to have two cameras, a microphone and a speaker. Humane crawled (or something like that) so Apple could walk. Apple will reportedly try to succeed where Humane failed (miserably) . On Wednesday, reported that the iPhone maker is working on an AI pin. The wearable is said to resemble a slightly thicker AirTag and include multiple cameras, a speaker, microphones, and wireless charging.


Grok's Leering Pictures Are the Newest Version of an Old Problem

Mother Jones

Grok's Leering Pictures Are the Newest Version of an Old Problem Image-based abuse predates Elon Musk's latest sleazy bot, but AI is making it worse. Get your news from a source that's not owned and controlled by oligarchs. There's a picture of myself that I had saved on my desktop for years; I suppose we could call it a caricature. A little more than a decade ago, someone on a Nazi messageboard pulled a photo of me from social media, then updated it with some antisemitic flair: a little cartoon rat sitting on my shoulder, a yellow pinned on its tiny body. Referencing what Jews were forced to wear during the Holocaust is meant to be a humiliation; the goal isn't hard to figure out, given that the whole star patch thing is near-medieval in both its imagery and its aims.