Education
Minimum Wasserstein distance estimator under covariate shift: closed-form, super-efficiency and irregularity
Lang, Junjun, Zhang, Qiong, Liu, Yukun
Covariate shift arises when covariate distributions differ between source and target populations while the conditional distribution of the response remains invariant, and it underlies problems in missing data and causal inference. We propose a minimum Wasserstein distance estimation framework for inference under covariate shift that avoids explicit modeling of outcome regressions or importance weights. The resulting W-estimator admits a closed-form expression and is numerically equivalent to the classical 1-nearest neighbor estimator, yielding a new optimal transport interpretation of nearest neighbor methods. We establish root-$n$ asymptotic normality and show that the estimator is not asymptotically linear, leading to super-efficiency relative to the semiparametric efficient estimator under covariate shift in certain regimes, and uniformly in missing data problems. Numerical simulations, along with an analysis of a rainfall dataset, underscore the exceptional performance of our W-estimator.
Optimal Transport under Group Fairness Constraints
Bleistein, Linus, Dagréou, Mathieu, Andrade, Francisco, Boudou, Thomas, Bellet, Aurélien
Ensuring fairness in matching algorithms is a key challenge in allocating scarce resources and positions. Focusing on Optimal Transport (OT), we introduce a novel notion of group fairness requiring that the probability of matching two individuals from any two given groups in the OT plan satisfies a predefined target. We first propose \texttt{FairSinkhorn}, a modified Sinkhorn algorithm to compute perfectly fair transport plans efficiently. Since exact fairness can significantly degrade matching quality in practice, we then develop two relaxation strategies. The first one involves solving a penalised OT problem, for which we derive novel finite-sample complexity guarantees. This result is of independent interest as it can be generalized to arbitrary convex penalties. Our second strategy leverages bilevel optimization to learn a ground cost that induces a fair OT solution, and we establish a bound guaranteeing that the learned cost yields fair matchings on unseen data. Finally, we present empirical results that illustrate the trade-offs between fairness and performance.
Match Made with Matrix Completion: Efficient Learning under Matching Interference
Tang, Zhiyuan, Chen, Wanning, Xu, Kan
Matching markets face increasing needs to learn the matching qualities between demand and supply for effective design of matching policies. In practice, the matching rewards are high-dimensional due to the growing diversity of participants. We leverage a natural low-rank matrix structure of the matching rewards in these two-sided markets, and propose to utilize matrix completion to accelerate reward learning with limited offline data. A unique property for matrix completion in this setting is that the entries of the reward matrix are observed with matching interference -- i.e., the entries are not observed independently but dependently due to matching or budget constraints. Such matching dependence renders unique technical challenges, such as sub-optimality or inapplicability of the existing analytical tools in the matrix completion literature, since they typically rely on sample independence. In this paper, we first show that standard nuclear norm regularization remains theoretically effective under matching interference. We provide a near-optimal Frobenius norm guarantee in this setting, coupled with a new analytical technique. Next, to guide certain matching decisions, we develop a novel ``double-enhanced'' estimator, based off the nuclear norm estimator, with a near-optimal entry-wise guarantee. Our double-enhancement procedure can apply to broader sampling schemes even with dependence, which may be of independent interest. Additionally, we extend our approach to online learning settings with matching constraints such as optimal matching and stable matching, and present improved regret bounds in matrix dimensions. Finally, we demonstrate the practical value of our methods using both synthetic data and real data of labor markets.
Implicit bias as a Gauge correction: Theory and Inverse Design
Aladrah, Nicola, Ballarin, Emanuele, Biagetti, Matteo, Ansuini, Alessio, d'Onofrio, Alberto, Anselmi, Fabio
A central problem in machine learning theory is to characterize how learning dynamics select particular solutions among the many compatible with the training objective, a phenomenon, called implicit bias, which remains only partially characterized. In the present work, we identify a general mechanism, in terms of an explicit geometric correction of the learning dynamics, for the emergence of implicit biases, arising from the interaction between continuous symmetries in the model's parametrization and stochasticity in the optimization process. Our viewpoint is constructive in two complementary directions: given model symmetries, one can derive the implicit bias they induce; conversely, one can inverse-design a wide class of different implicit biases by computing specific redundant parameterizations. More precisely, we show that, when the dynamics is expressed in the quotient space obtained by factoring out the symmetry group of the parameterization, the resulting stochastic differential equation gains a closed form geometric correction in the stationary distribution of the optimizer dynamics favoring orbits with small local volume. We compute the resulting symmetry induced bias for a range of architectures, showing how several well known results fit into a single unified framework. The approach also provides a practical methodology for deriving implicit biases in new settings, and it yields concrete, testable predictions that we confirm by numerical simulations on toy models trained on synthetic data, leaving more complex scenarios for future work. Finally, we test the implicit bias inverse-design procedure in notable cases, including biases toward sparsity in linear features or in spectral properties of the model parameters.
Parents of under-fives to be offered screen time guidance
Parents of under-fives in England are to be offered official advice on how long their children should spend watching TV or looking at computer screens. The government says it will publish its first guidance on screen time for the age group in April. It comes as government research was published showing that about 98% of children under two were watching screens on a daily basis - with parents, teachers and nursery staff saying youngsters were finding it harder to hold conversations or concentrate on learning. Children with the highest screen time - around five hours a day - reportedly could say significantly fewer words than those at the other end of the scale who watched for around 44 minutes. A national working group led by Children's Commissioner for England Dame Rachel de Souza and Department for Education scientific adviser Professor Russell Viner will formulate the guidance after speaking to parents, children and early years practitioners.
A brief note on learning problem with global perspectives
In this brief note, we considers the problem of learning with dynamic-optimizing principal-agent setting, in which the agents are allowed to have global perspectives about the learning process, i.e., the ability to view things according to their relative importances or in their true relations based-on some aggregated information shared by the principal. Whereas, the principal, which is exerting an influence on the learning process of the agents in the aggregation, is primarily tasked to solve a high-level optimization problem posed as an empirical-likelihood estimator under conditional moment restrictions model that also accounts information about the agents' predictive performances on out-of-samples as well as a set of private datasets available only to the principal (e.g., see [1], [2], [3], [4] and [5] for further discussions on empirical likelihood methods with moment restrictions). Here, we provide a coherent mathematical argument which is necessary for characterizing the learning process behind this abstract dynamic-optimizing principal-agent learning framework. Note that, due to the inherent feedbacks behavior among the agents, the proposed learning framework remarkably offers some advantages in terms of stability and consistency, despite that both the principal and the agents do not necessarily need to have any knowledge of the sample distributions or the quality of each others datasets. Finally, it is worth remarking that such a learning framework can provide new insights in the context of collaborative learning problem with global perspectives that exploits the principal-agent setting (e.g., see [6], [7], [8] or [9] for related discussions), although we acknowledge that there are a number of conceptual and theoretical problems, such as small sample properties, still need to be addressed.
From Unstructured Data to Demand Counterfactuals: Theory and Practice
Christensen, Timothy, Compiani, Giovanni
Empirical models of demand for differentiated products rely on low-dimensional product representations to capture substitution patterns. These representations are increasingly proxied by applying ML methods to high-dimensional, unstructured data, including product descriptions and images. When proxies fail to capture the true dimensions of differentiation that drive substitution, standard workflows will deliver biased counterfactuals and invalid inference. We develop a practical toolkit that corrects this bias and ensures valid inference for a broad class of counterfactuals. Our approach applies to market-level and/or individual data, requires minimal additional computation, is efficient, delivers simple formulas for standard errors, and accommodates data-dependent proxies, including embeddings from fine-tuned ML models. It can also be used with standard quantitative attributes when mismeasurement is a concern. In addition, we propose diagnostics to assess the adequacy of the proxy construction and dimension. The approach yields meaningful improvements in predicting counterfactual substitution in both simulations and an empirical application.
The Download: the case for AI slop, and helping CRISPR fulfill its promise
If I were to locate the moment AI slop broke through into popular consciousness, I'd pick the video of rabbits bouncing on a trampoline that went viral last summer. For many savvy internet users, myself included, it was the first time we were fooled by an AI video, and it ended up spawning a wave of almost identical generated clips. My first reaction was that, broadly speaking, all of this sucked. That's become a familiar refrain, in think pieces and at dinner parties. Everything online is slop now--the internet "enshittified," with AI taking much of the blame. But then friends started sharing AI clips in group chats that were compellingly weird, or funny.
America's new dietary guidelines ignore decades of scientific research
America's new dietary guidelines ignore decades of scientific research An emphasis on fruit, vegetables, and whole foods is welcome--but it's wrong to suggest steak and beef tallow should be prominent. The new year has barely begun, but the first days of 2026 have brought big news for health. On Monday, the US's federal health agency upended its recommendations for routine childhood vaccinations--a move that health associations worry puts children at unnecessary risk of preventable disease. There was more news from the federal government on Wednesday, when health secretary Robert F. Kennedy Jr. and his colleagues at the Departments of Health and Human Services and Agriculture unveiled new dietary guidelines for Americans . And they are causing a bit of a stir. RFK Jr's plan to improve America's diet is missing the point That's partly because they recommend products like red meat, butter, and beef tallow--foods that have been linked to cardiovascular disease, and that nutrition experts have been recommending people in their diets.
Learning Shrinks the Hard Tail: Training-Dependent Inference Scaling in a Solvable Linear Model
We analyze neural scaling laws in a solvable model of last-layer fine-tuning where targets have intrinsic, instance-heterogeneous difficulty. In our Latent Instance Difficulty (LID) model, each input's target variance is governed by a latent ``precision'' drawn from a heavy-tailed distribution. While generalization loss recovers standard scaling laws, our main contribution connects this to inference. The pass@$k$ failure rate exhibits a power-law decay, $k^{-β_\text{eff}}$, but the observed exponent $β_\text{eff}$ is training-dependent. It grows with sample size $N$ before saturating at an intrinsic limit $β$ set by the difficulty distribution's tail. This coupling reveals that learning shrinks the ``hard tail'' of the error distribution: improvements in the model's generalization error steepen the pass@$k$ curve until irreducible target variance dominates. The LID model yields testable, closed-form predictions for this behavior, including a compute-allocation rule that favors training before saturation and inference attempts after. We validate these predictions in simulations and in two real-data proxies: CIFAR-10H (human-label variance) and a maths teacher-student distillation task.