Europe
Briefly Noted Book Reviews
"The Lost Soldiers," "Homebound," "Once Upon a Time There Was Truth," and "My World Is Melting." The year is 1919, the midst of Bolshevik takeover in Ukraine, and twenty-eight Red Army soldiers have vanished into thin air, last seen at a bathhouse. Kolechko must track them down. He gets little help from the absurd locals, who range from obstinately useless to selfishly malicious. Kolechko is a kind of anti-Poirot--a fairly conventional man whose powers of detection lie not in a dazzling intuition but in a supernatural severed ear, which has a bug-like ability to pick up dialogue.
Generalization in Deep Neural Networks: Minimax Rates for Gradient Methods
Zhou, Junyu, Wang, Puyu, Lei, Yunwen, Kloft, Marius, Ying, Yiming
A central mystery in deep learning is how neural networks, despite being highly non-convex and heavily overparameterized, are able to achieve near-zero training error while still generalizing well to unseen data. This paradox has sparked a surge of research aimed at understanding the convergence and generalization behavior of neural networks [1, 2, 6, 7, 15, 38, 41, 49]. The Neural Tangent Kernel (NTK), introduced by [20], has become one of a foundational tool for understanding the behavior of training dynamics for neural networks, especially those trained using gradient-based methods such as gradient descent (GD) and stochastic gradient descent (SGD). The core idea here is to linearize the neural network around its random initialization, which enables the evolution of the network during training to be closely approximated by a kernel method associated with the corresponding NTK. This framework establishes a powerful connection between the evolution of a neural network during training process and the behavior of kernel methods in a reproducing kernel Hilbert space (RKHS) induced by the NTK, allowing insights from the kernel methods to inform our understanding of neural networks. Following this perspective, the influential work [34] showed that for regression problems, shallow neural networks trained by SGD can achieve generalization performance on par with their kernel counterparts.
Optimal Rates for Generalization of Gradient Descent Methods with Deep Neural Networks
Zhou, Junyu, Wang, Puyu, Lei, Yunwen, Ying, Yiming, Zhou, Ding-Xuan
Recent progress has been made in understanding the statistical generalization performance of gradient descent methods for overparameterized neural networks within the neural tangent kernel (NTK) regime. However, most of the existing work on regression problems is limited to shallow network architectures, leaving a notable gap in the theory of deep neural networks. This paper addresses this gap by presenting a comprehensive generalization analysis for deep ReLU networks trained using gradient descent (GD) and stochastic gradient descent (SGD). Specifically, we establish the first known minimax-optimal rates of excess population risk for both GD and SGD with deep ReLU networks, under the assumption that the network width scales polynomially with respect to the network depth and training sample size. Our results demonstrate that with sufficient width, gradient descent methods for deep ReLU networks can achieve optimal generalization rates on par with kernel methods.
Gen Z are refusing to buy rounds at the pub to avoid hangovers - now scientists say it really works
Caitlyn Jenner biographer and Robin Riker's ex William Hasley found dead on hiking trail at 78 Disgraceful texts'hot' teacher sent boy, 17, who she had illegal sex with where she moaned about her HUSBAND Everyone always said I cleared my throat a lot. But then I developed shoulder pain and doctors discovered the sinister cause... the world's deadliest cancer. Don't leave it too late like I did Urgent recall for 1.1m vehicles over fears they could spontaneously CATCH FIRE even when parked Moment Real Housewives star Lenny Hochstein's sexual assault accuser'dances' as she leaves Star Island mansion - before filing $100k civil lawsuit Leaked transcript of UNAIRED 60 Minutes interview exposes REAL reason'callous' CBS star Scott Pelley'deserved to be fired' Disturbing new death scene photos show tech whistleblower's haunting final moments... as forensic report casts doubt on suicide claims: 'Execution angle' 'Great' mom, 32, tried to gas herself and her three young kids to death after inviting them to'popcorn sleepover' in car, prosecutors allege The porn-fuelled fantasy middle-class husbands are desperate to try with their wives... and it almost always ends in divorce: JANA HOCKING The historic steel mill that helped build America was written off for dead. Medical student, 24, died by suicide in his white coat a day after he was suspended for alleged'inappropriate' behavior towards female patient, lawsuit alleges, as his heartbreaking goodbye note to parents is revealed John Oliver's private panic: Late-night curse spreads and host prepares for worst as insiders reveal his desperate'plan B'... and the industry whispers swirling about his fate Woke Vegas school compared boy to racist cross burner over pro-ICE stickers and expelled him... but did not punish pro-migrant students for class walkout, lawsuit alleges Gaming influencer Alex Cimo dies'very suddenly' aged 32 just a month after'refusing to accept his fate' Mother's final words before she was shot dead'by new husband' in front of her two young children All the backstage gossip from Miami Swim Week: Insider exposes'catty' VIP's diva demands... STEALING... and'morbidly embarrassing' celeb moment everyone is whispering about READ MORE: Gen Z are'zebra striping' on nights out to avoid hangovers From drinking'Tiger's milk' to soaking socks in vodka, many booze-loving Brits will try just about anything to avoid a hangover. Now, a new anti-hangover method is emerging on social media - avoiding rounds at the pub.
Painting bought for 100 in US charity shop sells for 190,000
A painting bought for less than $100 (£75) in a US charity shop in the 1960s has sold for almost £190,000 at auction. Art teacher Helene Plotkin bought the work by Scottish Colourist FCB Cadell in White Plains, New York in 1966, unaware of its true value. The painting, Interior: The Lady in Black, hung in her living room for 60 years - but the artist's signature was illegible and was only recently identified. It sold for £189,200, including buyer's premium, in Edinburgh as part of Lyon & Turnbull's Scottish painting and sculpture auction. The background to the painting only became clear when Helene's son Barry began his own research into it and took it for a valuation last year.
Adaptive Learning Rates with Surrogate Probability for Follow-the-Perturbed-Leader
Lee, Jongyeong, Honda, Junya, Ito, Shinji, Kim, Chansoo
Follow-the-regularized-leader framework has shown effectiveness and flexibility in online learning problems, where the choice of learning rates are known to be crucial. Recently, adaptive learning rates defined in terms of the arm-selection probabilities, obtained by solving convex optimization, have achieved improved best-of-both-worlds (BOBW) guarantees in various bandit problems. In contrast, BOBW guarantees for its computationally efficient alternative, follow-the-perturbed-leader (FTPL), remain relatively limited since its optimization-free nature ironically makes the design of adaptive, probability-dependent learning rates non-trivial. To address this challenge, we propose an adaptive learning rate for FTPL by introducing surrogate probability functions that can be computed only from the available quantities, without requiring the exact probabilities. Based on these learning rates with surrogate functions, we provide the BOBW guarantee for FTPL with Pareto perturbations for any shape parameter $α>1$, generalizing prior results restricted to specific choices of $α=2$. We further show the BOBW guarantees for FTPL with adaptive learning rates in the bandit problem with expert advices. Our approach preserves the computational simplicity of FTPL while enabling probability-dependent adaptivity, and the surrogate-based methodology may be of independent interest in other algorithmic frameworks beyond FTPL and learning rate designs.
Estimation of the sub-Gaussian parameter
Liu, Jason, Xu, Min, Xing, Jinchuan
The sub-Gaussian parameter (also called the variance proxy) of a mean-zero random variable $X$ is defined as $ξ^2_* = \sup_{λ\in \mathbb{R}} L(λ)$ where $L(λ) = \frac{2}{λ^2} \log \mathbb{E} e^{λX}$ is a weighted cumulant generating function. Despite the ubiquity of sub-Gaussian random variables, the estimation of $ξ^2_*$ has received little attention and is not yet well understood. In this work, we study a natural estimator of $ξ^2_*$ based on constrained maximization of the empirical analogue of $L$. We prove that the estimator is consistent bound the rates of convergence under assumptions on $L$: if $L$ has an maximizer, then our bound is $O_p(n^{-1/2 + \varepsilon})$ for any $\varepsilon > 0$; if the argmax of $L$ is also bounded, then the bound improves to $O_p(n^{-1/2})$. We show that our assumptions on $L$ are necessary by proving that the minimax risk over all sub-Gaussian distributions is $Ω(1)$; imposing increasingly strong assumptions on the tail growth of $L$ yields a continuum of classes whose minimax lower bound interpolates between $Ω(1/\log n)$ and $Ω(1)$. Root-n rate is possible if we restrict to a subclass of distributions where $L$ attains its supremum in a bounded region, in which case our estimator is minimax optimal. If the underlying distribution is not sub-Gaussian, we show that our estimator goes to infinity with a divergence rate controlled by the tail of the distribution. Finally, we apply our estimator in a Gene Ontology (GO) enrichment study to construct p-values for a large-scale permutation test, showing that it can serve as a reliable alternative to the peaks-over-threshold approach, particularly in regimes where the peaks-over-threshold method is of uncertain validity.
Conformal Risk Sharing: Certified Cost Allocation with Participation Guarantees
Sharing the financial impact of rare adverse events across a group can soften extreme individual burdens, but any participant made worse off by the arrangement has reason to leave. A credible mechanism must therefore provide each agent with a trustworthy cap on their future obligation and should be deployed only if the aggregate harm across participants is bounded. We formalise this as the Certified Allocation Problem: from finite data and without distributional assumptions, find a redistribution rule, produce obligation caps for every participant, and verify that no participant is made materially worse off. We propose Conformal Risk Sharing, which solves this problem by pairing an interpretable sharing policy with split conformal calibration. The sharing intensity is tuned on training data, while held-out calibration data produces distribution-free per-agent guarantees (valid under exchangeability). Experiments on synthetic and real-world data, including precipitation and energy-cooperative data, confirm that the framework can substantially reduce extreme obligations for high-risk agents while controlling harm to others.
Anchor PCA
Seiter, Benedikt, Fries, Anya, von Kügelgen, Julius, Peters, Jonas
Principal component analysis (PCA) is one of the most widely used unsupervised dimension reduction techniques. We study PCA for data from multiple related domains. Since principal components generally differ across domains, one way to obtain a shared low-rank embedding is to perform PCA on the pooled data. However, this approach can focus on spurious directions that exhibit high variation in only a few domains. To find a robust embedding that still explains most variance in unseen but similar domains, we propose instead to focus on shared directions of variation. To this end, we introduce Anchor PCA which trades off overall explained variance with agreement between the shared and domain-specific low-rank embeddings. Anchor PCA amounts to PCA on a modified target matrix and thus can be solved efficiently. Moreover, we show that Anchor PCA recovers a maximal invariant subspace and admits a minimax reconstruction interpretation under bounded domain-specific covariance inflations. On simulated and real-world gas sensor data with temporal drift, we demonstrate, respectively, that Anchor PCA recovers the maximally invariant subspace and yields embeddings that explain more variance on unseen domains than the pooling baseline and a worst-case alternative. Taken together, these findings establish Anchor PCA as a promising approach to robust unsupervised dimension reduction from multi-domain data.
My year with the robots: how Joanna Stern let AI into her home, work – and heart
In 2025, the tech journalist invited artificial intelligence to do nearly everything for her, including editing the book she was writing about the experiment. F or a year, Joanna Stern decided to turn herself into a "lab rat" - the object of her own experiment. Throughout 2025, she invited artificial intelligence into "every corner" of her life. She let AI answer her texts, decide what she ate and cooked, mow her lawn, fold her washing, drive her places, parse her mammograms and even, in the darkness of a burner phone, be her lover. The resulting book, I Am Not a Robot: My Year Using AI to Do (Almost) Everything, asks all the big questions, including: what happens when AI can do everything humans can do? And what comes after that?