Goto

Collaborating Authors

 Country


Sharpness Minimization Algorithms Do Not Only Minimize Sharpness To Achieve Better Generalization

Neural Information Processing Systems

Despite extensive studies, the underlying reason as to why overparameterizedneural networks can generalize remains elusive. Existing theory shows that common stochastic optimizers prefer flatter minimizers of the training loss, and thusa natural potential explanation is that flatness implies generalization. This workcritically examines this explanation. Through theoretical and empirical investigation, we identify the following three scenarios for two-layer ReLU networks: (1)flatness provably implies generalization; (2) there exist non-generalizing flattestmodels and sharpness minimization algorithms fail to generalize poorly, and (3)perhaps most strikingly, there exist non-generalizing flattest models, but sharpnessminimization algorithms still generalize. Our results suggest that the relationshipbetween sharpness and generalization subtly depends on the data distributionsand the model architectures and sharpness minimization algorithms do not onlyminimize sharpness to achieve better generalization.


Towards robust vision by multi-task learning on monkey visual cortex

Neural Information Processing Systems

Deep neural networks set the state-of-the-art across many tasks in computer vision, but their generalization ability to simple image distortions is surprisingly fragile. In contrast, the mammalian visual system is robust to a wide range of perturbations. Recent work suggests that this generalization ability can be explained by useful inductive biases encoded in the representations of visual stimuli throughout the visual cortex. Here, we successfully leveraged these inductive biases with a multi-task learning approach: we jointly trained a deep network to perform image classification and to predict neural activity in macaque primary visual cortex (V1) in response to the same natural stimuli. We measured the out-of-distribution generalization abilities of our resulting network by testing its robustness to common image distortions.


Elon Musk, AI and the antichrist: the biggest tech stories of 2025

The Guardian

Elon Musk receives a golden key from Donald Trump in the Oval Office at the White House in Washington DC on 30 May 2025. Elon Musk receives a golden key from Donald Trump in the Oval Office at the White House in Washington DC on 30 May 2025. I myself have a cold. Today, we are looking back at the biggest stories in tech of 2025 - Elon Musk's political rise, burst, and fall; artificial intelligence's subsumption of the global economy, all other technology, and even the Earth's topography; Australia's remarkable social media ban; the tech industry's new Trumpian politics; and, as a treat, a glimpse of the apocalypse offered by one of Silicon Valley's savviest and strangest billionaires. Tesla CEO Elon Musk attends a memorial service for slain far-right commentator Charlie Kirk at State Farm Stadium, in Glendale, Arizona, on 21 September 2025.


Smoothing the Landscape Boosts the Signal for SGD: Optimal Sample Complexity for Learning Single Index Models

Neural Information Processing Systems

We focus on the task of learning a single index model $\sigma(w^\star \cdot x)$ with respect to the isotropic Gaussian distribution in $d$ dimensions. Prior work has shown that the sample complexity of learning $w^\star$ is governed by the information exponent $k^\star$ of the link function $\sigma$, which is defined as the index of the first nonzero Hermite coefficient of $\sigma$. Ben Arous et al. (2021) showed that $n \gtrsim d^{k^\star-1}$ samples suffice for learning $w^\star$ and that this is tight for online SGD. However, the CSQ lower bound for gradient based methods only shows that $n \gtrsim d^{k^\star/2}$ samples are necessary. In this work, we close the gap between the upper and lower bounds by showing that online SGD on a smoothed loss learns $w^\star$ with $n \gtrsim d^{k^\star/2}$ samples. We also draw connections to statistical analyses of tensor PCA and to the implicit regularization effects of minibatch SGD on empirical losses.


Learning Composable Energy Surrogates for PDE Order Reduction

Neural Information Processing Systems

Meta-materials are an important emerging class of engineered materials in which complex macroscopic behaviour--whether electromagnetic, thermal, or mechanical--arises from modular substructure. Simulation and optimization of these materials are computationally challenging, as rich substructures necessitate high-fidelity finite element meshes to solve the governing PDEs. To address this, we leverage parametric modular structure to learn component-level surrogates, enabling cheaper high-fidelity simulation. We use a neural network to model the stored potential energy in a component given boundary conditions. This yields a structured prediction task: macroscopic behavior is determined by the minimizer of the system's total potential energy, which can be approximated by composing these surrogate models. Composable energy surrogates thus permit simulation in the reduced basis of component boundaries. Costly ground-truth simulation of the full structure is avoided, as training data are generated by performing finite element analysis of individual components. Using dataset aggregation to choose training data allows us to learn energy surrogates which produce accurate macroscopic behavior when composed, accelerating simulation of parametric meta-materials.


22 breathtaking images from the 2025 Landscape Photographer of the Year awards

Popular Science

Breakthroughs, discoveries, and DIY tips sent every weekday. From Iceland's spectacular fire and ice landscapes to Yemen's otherworldly Socotra dragon trees, our home planet hosts a diverse lineup of jaw-dropping scenery. The 12th annual International Landscape Photographer of the Year award honor professional and amateur photographers who venture far and wide to capture nature's beauty. Why do we have five fingers and toes? Breakthroughs, discoveries, and DIY tips sent every weekday.


Minimax Classification with 0-1 Loss and Performance Guarantees

Neural Information Processing Systems

Supervised classification techniques use training samples to find classification rules with small expected 0-1 loss. Conventional methods achieve efficient learning and out-of-sample generalization by minimizing surrogate losses over specific families of rules. This paper presents minimax risk classifiers (MRCs) that do not rely on a choice of surrogate loss and family of rules. MRCs achieve efficient learning and out-of-sample generalization by minimizing worst-case expected 0-1 loss w.r.t.


On-Demand Sampling: Learning Optimally from Multiple Distributions

Neural Information Processing Systems

Societal and real-world considerations such as robustness, fairness, social welfare and multi-agent tradeoffs have given rise to multi-distribution learning paradigms, such as collaborative [Blum et al. 2017], group distributionally robust [Sagawa et al. 2019], and fair federated learning [Mohri et al. 2019]. In each of these settings, a learner seeks to minimize its worstcase loss over a set of $n$ predefined distributions, while using as few samples as possible. In this paper, we establish the optimal sample complexity of these learning paradigms and give algorithms that meet this sample complexity. Importantly, our sample complexity bounds exceed that of the sample complexity of learning a single distribution only by an additive factor of $\frac{n\log(n)}{\epsilon^2}$. These improve upon the best known sample complexity of agnostic federated learning by Mohri et al. 2019 by a multiplicative factor of $n$, the sample complexity of collaborative learning by Nguyen and Zakynthinou 2018 by a multiplicative factor $\frac{\log(n)}{\epsilon^3}$, and give the first sample complexity bounds for the group DRO objective of Sagawa et al. 2019. To achieve optimal sample complexity, our algorithms learn to sample and learn from distributions on demand. Our algorithm design and analysis extends stochastic optimization techniques to solve zero-sum games in a new stochastic setting.


Cozy up (safely) to an e-scooter's lithium battery yule log

Popular Science

Breakthroughs, discoveries, and DIY tips sent every weekday. The United States Consumer Product Safety Commission (CPSC) is well known for getting their point across on social media. A seven-minute montage of mannequins succumbing to 4th of July firework injuries may be an unconventional way to warn about the dangers of recreational explosives--but try forgetting those images when lighting your next bottle rocket. In similar pyrotechnic fashion, the CPSC is warning everyone to take extra care during the holidays when it comes to all kinds of combustible, seasonally appropriate objects. On December 22, the commission illustrated how some gifts are far more flammable than others with its 30-minute Escooter Lithium-Ion Battery Yule Log video.


PC prices could rise by 8% in 2026 due to memory shortages

PCWorld

When you purchase through links in our articles, we may earn a small commission. At the same time, IDC predicts that the PC market could also shrink by 8.9 per cent during the year. Research firm IDC predicts that the average price of computers could rise by up to 8% in 2026 due to a global shortage of memory chips (RAM and NAND). The background is the sharp increase in demand for HBM memory for AI data centers, which is prioritized over the production of consumer memory, which is less profitable for manufacturers. Meanwhile, IDC predicts that the PC market could also shrink by between 2.4 and 8.9 percent in 2026 as a result of the shortage.