Goto

Collaborating Authors

 Industry


Efficient Swap Multicalibration of Elicitable Properties

arXiv.org Machine Learning

Multicalibration [HJKRR18] is an algorithmic fairness perspective that demands that the predictions of a predictor are correct conditional on themselves and membership in a collection of potentially overlapping subgroups of a population. The work of [NR23] established a surprising connection between multicalibration for an arbitrary property $ฮ“$ (e.g., mean or median) and property elicitation: a property $ฮ“$ can be multicalibrated if and only if it is elicitable, where elicitability is the notion that the true property value of a distribution can be obtained by solving a regression problem over the distribution. In the online setting, [NR23] proposed an inefficient algorithm that achieves $\sqrt T$ $\ell_2$-multicalibration error for a hypothesis class of group membership functions and an elicitable property $ฮ“$, after $T$ rounds of interaction between a forecaster and adversary. In this paper, we generalize multicalibration for an elicitable property $ฮ“$ from group membership functions to arbitrary bounded hypothesis classes and introduce a stronger notion -- swap multicalibration, following [GKR23]. Subsequently, we propose an oracle-efficient algorithm which, when given access to an online agnostic learner, achieves $T^{1/(r+1)}$ $\ell_r$-swap multicalibration error with high probability (for $r\ge2$) for a hypothesis class with bounded sequential Rademacher complexity and an elicitable property $ฮ“$. For the special case of $r=2$, this implies an oracle-efficient algorithm that achieves $T^{1/3}$ $\ell_2$-swap multicalibration error, which significantly improves on the previously established bounds for the problem [NR23, GMS25, LSS25a], and completely resolves an open question raised in [GJRR24] on the possibility of an oracle-efficient algorithm that achieves $\sqrt{T}$ $\ell_2$-mean multicalibration error by answering it in a strongly affirmative sense.


Trained on Tokens, Calibrated on Concepts: The Emergence of Semantic Calibration in LLMs

arXiv.org Machine Learning

Large Language Models (LLMs) often lack meaningful confidence estimates for their outputs. While base LLMs are known to exhibit next-token calibration, it remains unclear whether they can assess confidence in the actual meaning of their responses beyond the token level. We find that, when using a certain sampling-based notion of semantic calibration, base LLMs are remarkably well-calibrated: they can meaningfully assess confidence in open-domain question-answering tasks, despite not being explicitly trained to do so. Our main theoretical contribution establishes a mechanism for why semantic calibration emerges as a byproduct of next-token prediction, leveraging a recent connection between calibration and local loss optimality. The theory relies on a general definition of "B-calibration," which is a notion of calibration parameterized by a choice of equivalence classes (semantic or otherwise). This theoretical mechanism leads to a testable prediction: base LLMs will be semantically calibrated when they can easily predict their own distribution over semantic answer classes before generating a response. We state three implications of this prediction, which we validate through experiments: (1) Base LLMs are semantically calibrated across question-answering tasks, (2) RL instruction-tuning systematically breaks this calibration, and (3) chain-of-thought reasoning breaks calibration. To our knowledge, our work provides the first principled explanation of when and why semantic calibration emerges in LLMs.


Provable Separations between Memorization and Generalization in Diffusion Models

arXiv.org Machine Learning

Diffusion models have achieved remarkable success across diverse domains, but they remain vulnerable to memorization -- reproducing training data rather than generating novel outputs. This not only limits their creative potential but also raises concerns about privacy and safety. While empirical studies have explored mitigation strategies, theoretical understanding of memorization remains limited. We address this gap through developing a dual-separation result via two complementary perspectives: statistical estimation and network approximation. From the estimation side, we show that the ground-truth score function does not minimize the empirical denoising loss, creating a separation that drives memorization. From the approximation side, we prove that implementing the empirical score function requires network size to scale with sample size, spelling a separation compared to the more compact network representation of the ground-truth score function. Guided by these insights, we develop a pruning-based method that reduces memorization while maintaining generation quality in diffusion transformers.


Calibrated Principal Component Regression

arXiv.org Machine Learning

We propose a new method for statistical inference in generalized linear models. In the overparameterized regime, Principal Component Regression (PCR) reduces variance by projecting high-dimensional data to a low-dimensional principal subspace before fitting. However, PCR incurs truncation bias whenever the true regression vector has mass outside the retained principal components (PC). To mitigate the bias, we propose Calibrated Principal Component Regression (CPCR), which first learns a low-variance prior in the PC subspace and then calibrates the model in the original feature space via a centered Tikhonov step. CPCR leverages cross-fitting and controls the truncation bias by softening PCR's hard cutoff. Theoretically, we calculate the out-of-sample risk in the random matrix regime, which shows that CPCR outperforms standard PCR when the regression signal has non-negligible components in low-variance directions. Empirically, CPCR consistently improves prediction across multiple overparameterized problems. The results highlight CPCR's stability and flexibility in modern overparameterized settings.


Ukraine drone strikes throw power supplies into disarray in Russian cities

Al Jazeera

Is Trump losing patience with Putin? Will sanctions against Russian oil giants hurt Putin? Ukraine has hit back at Russia's attempts to disable its energy infrastructure with air strikes that succeeded in disrupting power and heating in two cities across the border. Alexander Gusev, regional governor of Voronezh, said several drones were electronically jammed over the city - home to more than one million people - and sparked a fire at a local utility facility that was quickly extinguished. A Russian Defence Ministry statement made no mention of either the Voronezh or Belgorod areas, reporting 44 Ukrainian drones were destroyed or intercepted by Russian forces during the night.


Thieves steal 100M in jewels from Louvre after museum used own name as surveillance password

FOX News

Thieves stole $100 million in jewels from the Louvre Museum in Paris after exploiting weak passwords, including using "Louvre" as a surveillance system password.


The State of AI: Energy is king, and the US is falling behind

MIT Technology Review

This week, Casey Crownhart, senior reporter for energy at MIT Technology Review and Pilita Clark, FT's columnist, consider how China's rapid renewables buildout could help it leapfrog on AI progress. In the age of AI, the biggest barrier to progress isn't money but energy . That should be particularly worrying here in the US, where massive data centers are waiting to come online, and it doesn't look as if the country will build the steady power supply or infrastructure needed to serve them all. For about a decade before 2020, data centers were able to offset increased demand with efficiency improvements . Now, though, electricity demand is ticking up in the US, with billions of queries to popular AI models each day--and efficiency gains aren't keeping pace. With too little new power capacity coming online, the strain is starting to show: Electricity bills are ballooning for people who live in places where data centers place a growing load on the grid.


13 perfect panoramic images from the 2025 Epson International Pano awards

Popular Science

Taken in Algeria, 'Last Fireworks' is this year's first place category and open competition overall winner. Breakthroughs, discoveries, and DIY tips sent every weekday. The winners of the 2025 Epson International Pano awards have been announced, showcasing photographs of our great, big, beautiful world in ultra-wide glory. Italy's Alex Wides (Alessandro Cantarelli) won the Open Photographer of the Year and the Nature/Landscape category for his fine-art landscapes (seen above and below). Among this year's 3,423 entries, there were more photographs of the Northern Lights than usual, coinciding with the 11-year solar cycle maximum .


AI could drive US unemployment to 20%, senators warn as new bill targets job tracking

FOX News

Senator Josh Hawley and Senator Mark Warner introduce the AI-Related Job Impacts Clarity Act requiring companies to report AI-related job impacts to the Department of Labor.


I'm a committed introvert โ€“ but no AI will take away the joy I get from other people Emma Beddington

The Guardian

'I'm baffled how anyone could use AI to participate in a hobby.' 'I'm baffled how anyone could use AI to participate in a hobby.' I'm a committed introvert - but no AI will take away the joy I get from other people T his is depressing: according to the Cut, people are using AI to solve escape room puzzles and cheat at trivia nights. Surely, that is the definition of spoiling your own fun? "Like going into a corn maze and just wanting a straight line to the end," says one TikToker quoted in the article. There's also an interview with a keen reader who uses ChatGPT as a book club replacement, scraping the internet and aggregating "stimulating opinions and perspectives". All well and good (actually, no, it sounds bleak as hell) until he had a character's death spoilered in the fantasy epic he had been enjoying.