Goto

Collaborating Authors

 Education


Dyslexia and the Reading Wars

The New Yorker

Proven methods for teaching the readers who struggle most have been known for decades. Why do we often fail to use them? "There's a window of opportunity to intervene," Mark Seidenberg, a cognitive neuroscientist, said. "You don't want to let that go." In 2024, my niece Caroline received a Ph.D. in gravitational-wave physics. Her research interests include "the impact of model inaccuracies on biases in parameters recovered from gravitational wave data" and "Petrov type, principal null directions, and Killing tensors of slowly rotating black holes in quadratic gravity." I watched a little of her dissertation defense, on Zoom, and was lost as soon as she'd finished introducing herself. She and her husband now live in Italy, where she has a postdoctoral appointment. Caroline's academic achievements seem especially impressive if you know that until third grade she could barely read: to her, words on a page looked like a pulsing mass. She attended a private school in Connecticut, and there was a set time every day when students selected books to read on their own. "I can't remember how long that lasted, but it felt endless," she told me. She hid her disability by turning pages when her classmates did, and by volunteering to draw illustrations during group story-writing projects. One day, she told her grandmother that she could sound out individual letters but when she got to "the end of a row" she couldn't remember what had come before. A psychologist eventually identified her condition as dyslexia. Fluent readers sometimes think of dyslexia as a tendency to put letters in the wrong order or facing the wrong direction, but it's more complicated than that.


Towards Sharp Minimax Risk Bounds for Operator Learning

arXiv.org Machine Learning

A new paradigm in machine learning for scientific computing is focused on designing learning algorithms and methods for continuum problems. This paradigm is referred to as operator learning and has received considerable interest in the last few years [5,7,18,20,23-25,27,30,34,36]. The basic task may be posed as learning a map between infinite-dimensional function spaces, i.e., learning an operator F: X Y, where, for example, X and Y are real, separable Hilbert spaces. Operator learning naturally arises in many scientific problems where one wants to learn how a continuum model, often described by partial differential equations (PDEs), maps inputs, such as parameters or boundary conditions, to outputs, such as states or observables. A prototypical example to keep in mind is learning parameter-to-solution maps of parametric PDEs [1,2,11]. In contrast to more classical surrogate modeling, which typically focuses on learning finite-dimensional parameter-to-solution maps for some fixed discretization, operator learning directly aims to learn/approximate the continuum map F: X Y itself. Thus, the inputs and outputs are functions (not vectors) and the goal is to directly design discretization-invariant methods [7,23]. From a statistical perspective, this naturally leads to a nonparametric regression problem in which both the object of interest (the operator) and the observations (finite number of noisy samples) are infinite-dimensional.


Convergence Guarantees for Federated SARSA with Local Training and Heterogeneous Agents

arXiv.org Machine Learning

We present a novel theoretical analysis of Federated SARSA (FedSARSA) with linear function approximation and local training. We establish convergence guarantees for FedSARSA in the presence of heterogeneity, both in local transitions and rewards, providing the first sample and communication complexity bounds in this setting. At the core of our analysis is a new, exact multi-step error expansion for single-agent SARSA, which is of independent interest. Our analysis precisely quantifies the impact of heterogeneity, demonstrating the convergence of FedSARSA with multiple local updates. Crucially, we show that FedSARSA achieves linear speed-up with respect to the number of agents, up to higher-order terms due to Markovian sampling. Numerical experiments support our theoretical findings.


Research Reveals the Optimal Way to Optimize

WIRED

The leading approach to the simplex method, a widely used technique for balancing complex logistical constraints, can't get any better. In 1939, upon arriving late to his statistics course at UC Berkeley, George Dantzig--a first-year graduate student--copied two problems off the blackboard, thinking they were a homework assignment. He found the homework "harder to do than usual," he would later recount, and apologized to the professor for taking some extra days to complete it. A few weeks later, his professor told him that he had solved two famous open problems in statistics. Dantzig's work would provide the basis for his doctoral dissertation and, decades later, inspiration for the film .


Would You Trust a 22-Year-Old AI Billionaire With the Global Economy?

The Atlantic - Technology

B rendan Foody is 22 years old and runs a company worth billions. This August, I met the young CEO in a glass conference room overlooking the San Francisco Bay. While his peers are searching for their first jobs, Foody is pursuing a " master plan," as he calls it, to upend the global labor market. His start-up, Mercor, offers an AI-powered hiring platform: Bots weed through résumés, and even conduct interviews. In the next five years, Foody told me, AI could automate 50 percent of the tasks that people do today.


The 20 best video games of 2025

The Guardian

An arena warrior on a losing streak takes refuge in a vast forest where she discovers the joy of working in a cosy teashop. From this simple premise comes a joyful game of mindfulness and social interaction, as Alta learns how to serve up witty conversation and decent hot drinks. Colourful and highly stylised, it is a thoughtful study of burnout and recovery. An attempted-murder mystery set in an a 1920s all-girls private school reveals itself to also be an eviscerating takedown of British class politics. Witty and beautifully drawn, it is full of amusing boarding-school stereotypes, from self-interested prefects to a terrifying matron, whose motivations and personal grievances must be slowly unpicked.


Advantages and limitations in the use of transfer learning for individual treatment effects in causal machine learning

arXiv.org Machine Learning

Generalizing causal knowledge across diverse environments is challenging, especially when estimates from large-scale datasets must be applied to smaller or systematically different contexts, where external validity is critical. Model-based estimators of individual treatment effects (ITE) from machine learning require large sample sizes, limiting their applicability in domains such as behavioral sciences with smaller datasets. We demonstrate how estimation of ITEs with Treatment Agnostic Representation Networks (TARNet; Shalit et al., 2017) can be improved by leveraging knowledge from source datasets and adapting it to new settings via transfer learning (TL-TARNet; Aloui et al., 2023). In simulations that vary source and sample sizes and consider both randomized and non-randomized intervention target settings, the transfer-learning extension TL-TARNet improves upon standard TARNet, reducing ITE error and attenuating bias when a large unbiased source is available and target samples are small. In an empirical application using the India Human Development Survey (IHDS-II), we estimate the effect of mothers' firewood collection time on children's weekly study time; transfer learning pulls the target mean ITEs toward the source ITE estimate, reducing bias in the estimates obtained without transfer. These results suggest that transfer learning for causal models can improve the estimation of ITE in small samples.


2025 AAAI / ACM SIGAI Doctoral Consortium interviews compilation

AIHub

Authors pictured in order of their interview publication date (left to right, top to bottom). Each year, a small group of PhD students are chosen to participate in the AAAI/SIGAI Doctoral Consortium . This initiative provides an opportunity for the students to discuss and explore their research interests and career objectives in an interdisciplinary workshop together with a panel of established researchers. During 2025, we met with some of the students to find out more about their research and the doctoral consortium experience. Kunpeng Xu completed his PhD at the Université de Sherbrooke and is now a postdoctoral fellow at McGill University.


Over half of deepfakes of underage victims made by classmates, Japanese police say

The Japan Times

The National Police Agency plans to warn against the obscene use of AI at delinquency-prevention lectures at schools and other events. More than half of cases reported to Japanese police of explicit deepfakes targeting those aged under 18 were created with the involvement of students from the same schools as the victims, National Police Agency data have shown. This is the first time that the NPA has released information on minors who became victims of obscene fake images created using generative artificial intelligence and other technologies. The agency plans to create flyers and warn against such use of AI at delinquency-prevention lectures at schools and other locations. According to the NPA, police were consulted over 79 cases of deepfakes targeting those up to the age of 17 from January to September this year.


A Teacher-Student Perspective on the Dynamics of Learning Near the Optimal Point

arXiv.org Machine Learning

Near an optimal learning point of a neural network, the learning performance of gradient descent dynamics is dictated by the Hessian matrix of the loss function with respect to the network parameters. We characterize the Hessian eigenspectrum for some classes of teacher-student problems, when the teacher and student networks have matching weights, showing that the smaller eigenvalues of the Hessian determine long-time learning performance. For linear networks, we analytically establish that for large networks the spectrum asymptotically follows a convolution of a scaled chi-square distribution with a scaled Marchenko-Pastur distribution. We numerically analyse the Hessian spectrum for polynomial and other non-linear networks. Furthermore, we show that the rank of the Hessian matrix can be seen as an effective number of parameters for networks using polynomial activation functions. For a generic non-linear activation function, such as the error function, we empirically observe that the Hessian matrix is always full rank.