Deep Learning
Scaling LLM Planning: NL2FLOW for Parametric Problem Generation and Rigorous Evaluation
Robust workflow composition is critical for effective agent performance, yet progress in Large Language Model (LLM) planning and reasoning is hindered by a scarcity of scalable evaluation data. This work introduces NL2Flow, a fully automated pipeline for generating and evaluating workflow planning problems. NL2Flow generates problems parametrically in a structured intermediate representation, translating them into both natural language and formal PDDL. I evaluate several open-source, instruct-tuned LLMs on a dataset of 2296 low-difficulty problems generated by NL2Flow. Results demonstrate that the best-performing model achieved 86% success in generating valid plans and 69% in generating optimal plans (for solvable problems). Regression analysis shows that the influence of problem characteristics on plan generation is contingent on both model and prompt design. Importantly, translating natural language problems into a structured JSON representation prior to symbolic planning significantly improved success rates, suggesting a benefit from neuro-symbolic integration. These findings underscore the importance of understanding error sources within LLM reasoning as systems scale to more complex tasks. As LLM reasoning scales to increasingly complex problems, understanding the shifting bottlenecks and sources of error within these systems will be crucial.
PriorGuide: Test-Time Prior Adaptation for Simulation-Based Inference
Yang, Yang, Rissanen, Severi, Chang, Paul E., Loka, Nasrulloh, Huang, Daolang, Solin, Arno, Heinonen, Markus, Acerbi, Luigi
Amortized simulator-based inference offers a powerful framework for tackling Bayesian inference in computational fields such as engineering or neuroscience, increasingly leveraging modern generative methods like diffusion models to map observed data to model parameters or future predictions. These approaches yield posterior or posterior-predictive samples for new datasets without requiring further simulator calls after training on simulated parameter-data pairs. However, their applicability is often limited by the prior distribution(s) used to generate model parameters during this training phase. To overcome this constraint, we introduce PriorGuide, a technique specifically designed for diffusion-based amortized inference methods. PriorGuide leverages a novel guidance approximation that enables flexible adaptation of the trained diffusion model to new priors at test time, crucially without costly retraining. This allows users to readily incorporate updated information or expert knowledge post-training, enhancing the versatility of pre-trained inference models.
Conformal Inference for Open-Set and Imbalanced Classification
Xie, Tianmin, Zhou, Yanfei, Liang, Ziyi, Favaro, Stefano, Sesia, Matteo
This paper presents a conformal prediction method for classification in highly imbalanced and open-set settings, where there are many possible classes and not all may be represented in the data. Existing approaches require a finite, known label space and typically involve random sample splitting, which works well when there is a sufficient number of observations from each class. Consequently, they have two limitations: (i) they fail to provide adequate coverage when encountering new labels at test time, and (ii) they may become overly conservative when predicting previously seen labels. To obtain valid prediction sets in the presence of unseen labels, we compute and integrate into our predictions a new family of conformal p-values that can test whether a new data point belongs to a previously unseen class. We study these p-values theoretically, establishing their optimality, and uncover an intriguing connection with the classical Good--Turing estimator for the probability of observing a new species. To make more efficient use of imbalanced data, we also develop a selective sample splitting algorithm that partitions training and calibration data based on label frequency, leading to more informative predictions. Despite breaking exchangeability, this allows maintaining finite-sample guarantees through suitable re-weighting. With both simulated and real data, we demonstrate our method leads to prediction sets with valid coverage even in challenging open-set scenarios with infinite numbers of possible labels, and produces more informative predictions under extreme class imbalance.
A Connection Between Score Matching and Local Intrinsic Dimension
Yeats, Eric, Jacobson, Aaron, Hannan, Darryl, Jia, Yiran, Doster, Timothy, Kvinge, Henry, Mahan, Scott
The local intrinsic dimension (LID) of data is a fundamental quantity in signal processing and learning theory, but quantifying the LID of high-dimensional, complex data has been a historically challenging task. Recent works have discovered that diffusion models capture the LID of data through the spectra of their score estimates and through the rate of change of their density estimates under various noise perturbations. While these methods can accurately quantify LID, they require either many forward passes of the diffusion model or use of gradient computation, limiting their applicability in compute- and memory-constrained scenarios. We show that the LID is a lower bound on the denoising score matching loss, motivating use of the denoising score matching loss as a LID estimator. Moreover, we show that the equivalent implicit score matching loss also approximates LID via the normal dimension and is closely related to a recent LID estimator, FLIPD. Our experiments on a manifold benchmark and with Stable Diffusion 3.5 indicate that the denoising score matching loss is a highly competitive and scalable LID estimator, achieving superior accuracy and memory footprint under increasing problem size and quantization level.
Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning
Oh, Junsoo, Huang, Wei, Suzuki, Taiji
Mamba, a recently proposed linear-time sequence model, has attracted significant attention for its computational efficiency and strong empirical performance. However, a rigorous theoretical understanding of its underlying mechanisms remains limited. In this work, we provide a theoretical analysis of Mamba's in-context learning (ICL) capability by focusing on tasks defined by low-dimensional nonlinear target functions. Specifically, we study in-context learning of a single-index model $y \approx g_*(\langle \boldsymbolฮฒ, \boldsymbol{x} \rangle)$, which depends on only a single relevant direction $\boldsymbolฮฒ$, referred to as feature. We prove that Mamba, pretrained by gradient-based methods, can achieve efficient ICL via test-time feature learning, extracting the relevant direction directly from context examples. Consequently, we establish a test-time sample complexity that improves upon linear Transformers -- analyzed to behave like kernel methods -- and is comparable to nonlinear Transformers, which have been shown to surpass the Correlational Statistical Query (CSQ) lower bound and achieve near information-theoretically optimal rate in previous works. Our analysis reveals the crucial role of the nonlinear gating mechanism in Mamba for feature extraction, highlighting it as the fundamental driver behind Mamba's ability to achieve both computational efficiency and high performance.
The AI Industry's Scaling Obsession Is Headed for a Cliff
The AI Industry's Scaling Obsession Is Headed for a Cliff Huge AI infrastructure deals assume that algorithms will keep improving with scale. A new study from MIT suggests the biggest and most computationally intensive AI models may soon offer diminishing returns compared to smaller models. By mapping scaling laws against continued improvements in model efficiency, the researchers found that it could become harder to wring leaps in performance from giant models whereas efficiency gains could make models running on more modest hardware increasingly capable over the next decade. "In the next five to 10 years, things are very likely to start narrowing," says Neil Thompson, a computer scientist and professor at MIT involved in the study. Leaps in efficiency, like those seen with DeepSeek's remarkably low-cost model in January, have already served as a reality check for the AI industry, which is accustomed to burning massive amounts of compute.
Gemini for Home's daily briefings are getting spooky, users say
When you purchase through links in our articles, we may earn a small commission. Gemini for Home's daily briefings are getting spooky, users say Halloween decorations, among other things, are playing tricks on the daily smart home summaries generated by Google's Gemini, according to some users. Those are just some of the things that Google's Gemini have been reporting in its Home Briefs--the summaries it can produce of the daily goings-on detected by Nest security cameras and other connected smart home devices--and some Gemini for Home users say they're getting thoroughly creeped out by the briefings, particularly with Halloween right around the corner. "Throughout the morning, several instances of people in black cloaks or robes were observed standing in the yard," read a Home Brief screenshot posed by a Google Home user on Reddit . "The unusual presence of individuals in black cloaks or robes continued into the afternoon, with multiple sighting in the yard and approaching the driveway."
The AI bubble is heading towards a burst but it won't be the end of AI
The AI bubble is heading towards a burst but it won't be the end of AI Economists, bankers and even the boss of OpenAI are warning of a rapidly inflating AI bubble. If and when it bursts, what will happen to the technological breakthroughs of the past few years? The hundreds of billions of dollars being spent on AI seem to have inflated a global financial bubble that's now fit to burst, leaving companies and investors at risk of holding vast debt that cannot be serviced by the meagre revenue brought in by current AI services. But what does that mean for the future of the technology underpinning this financial feeding frenzy? In recent weeks, warnings of a potential AI bubble have come from the International Monetary Fund, the Bank of England, the head of the largest US bank, and even OpenAI boss Sam Altman .
The Download: Big Tech's carbon removals plans, and the next wave of nuclear reactors
Microsoft, JP MorganChase, and a tech company consortium that includes Alphabet, Meta, Shopify, and Stripe have all recently struck multimillion-dollar deals to pay paper mill owners to capture at least hundreds of thousands of tons of this greenhouse gas by installing carbon scrubbing equipment in their facilities. The captured carbon dioxide will then be piped down into saline aquifers more than a mile underground, where it should be sequestered permanently. Big Tech is suddenly betting big on this form of carbon removal, known as bioenergy with carbon capture and storage, or BECCS. But experts have raised a number of concerns. Like many new nuclear startups, Kairos promises a path to reliable, 24/7 decarbonized power. Unlike most, it already has prototypes under construction and permits for several reactors.