Industry
Coupled Training with Privileged Information and Unlabeled Data
Shi, Jiahao, Hagrass, Omar, Klusowski, Jason M.
In many prediction problems, we have extra information during training (for example, measurements that are expensive or slow to collect) that will not be available when the model is deployed. A common strategy is to first train a model that uses all training information, then use its predictions on unlabeled examples to train a second model that only uses the inputs available at test time. However, when the extra training-only information is weak or noisy, this Two-Stage approach can mislead the deployment model and even hurt accuracy. We propose a joint training method that learns the two models together, so the deployment model can benefit from the extra information only when it actually helps, instead of inheriting its mistakes. We provide guarantees that describe when joint training improves prediction accuracy and analyze a simple alternating training algorithm for large, high-dimensional models. Experiments on synthetic data and real-world prediction tasks show that our approach avoids these failures and robustly outperforms standard Two-Stage baselines.
When Is Next-Token Prediction Useful? Marginalization, Ergodicity, Mixture Identifiability, Local Sufficiency, RAG, Tools, and Programming
Language models trained on observed sequences are often described as learning the conditional distribution of the next token given previous tokens. This description is only conditionally correct. A model trained on realized token trajectories does not observe full conditional laws; it receives sampled continuations. Moreover, real language generation is conditioned not only on previous words but also on non-textual circumstances: facts, events, intentions, goals, beliefs, social context, and task-specific constraints. This paper distinguishes three objects that are often conflated: the full conditional language process conditioned on latent circumstances, the marginal text-only process obtained by integrating those circumstances out, and the model-induced distribution learned from finite observed corpora. The paper argues that interpreting model training as estimating the marginal text-only law requires strong assumptions of stationarity, representativeness, and ergodicity, assumptions that are standard in statistical estimation but problematic when applied to heterogeneous language corpora. Even if these assumptions hold, the marginal text-only law is useful only when the observed prefix is an approximately sufficient statistic for the latent circumstances relevant to continuation. In information-theoretic terms, usefulness requires that the residual conditional mutual information between the next token and the omitted circumstances, given the observed text, be small. The paper then extends this argument to heterogeneous training corpora. Finally, the paper interprets Retrieval Augmented Generation (RAG) and tool use as conditional sufficiency devices.
Concomitant DAG Learning: On the Roles of Noise Adaptivity, Sparsity, and Non-negativity
Mateos, Gonzalo, Rey, Samuel, Ajorlou, Hamed, Tepper, Mariano
Directed acyclic graphs (DAGs) constitute a central modeling tool to enable principled reasoning about cause-effect interactions in complex systems. However, since the causal structure underlying a group of variables is often unknown and interventions may be infeasible or ethically challenging to implement, there is a need to address the task of inferring DAGs from observational data. However, most classical structure identification approaches face two key obstacles: the combinatorial challenge of enforcing acyclicity, which severely limits scalability, and identifiability challenges arising from latent confounding or heterogeneous noise. This tutorial offers an overview of recent signal processing and optimization advances that address these issues by recasting DAG structure learning as a continuous, score-based estimation problem over adjacency matrices. We begin with a didactic introduction to structural equation models and the formulation of causal graph recovery, followed by a historical survey of score-based methods ranging from early combinatorial search schemes and greedy heuristics to modern continuous frameworks that leverage smooth characterizations of acyclicity. Building on this foundation, we describe concomitant DAG estimation methods that jointly infer sparse causal structure and exogenous noise levels, improving robustness under heteroscedasticity and distribution shifts by rendering the estimator noise adaptive. All in all, the tutorial introduces readers to challenges and opportunities for signal processing research at the crossroads of causal inference, high-dimensional statistics, and scalable graph learning, while outlining emerging directions including online, nonlinear, and neural causal discovery.
Asymmetric Scaling Laws from Sparse Features
We introduce a model for neural scaling laws under sparse activations. In the model, test loss is often dominated by rare coordinates that are never observed in the training input. This mechanism induces a novel bottleneck absent from dense models. We derive the asymptotic population loss in both the underparameterized and overparameterized regimes, and show that the loss exhibits a double-descent peak near the interpolation threshold -- where the number of parameters is just sufficient to fit the training data -- resulting in a loss curve governed by two distinct scaling exponents -- one for the overparameterized regime and one for the underparameterized regime -- with a gap determined by the degree of sparsity. Additionally, we derive a compute-optimal frontier that favors increasing dataset size over model capacity under fixed compute budgets. We also analyze gradient-descent dynamics and identify a scaling law for the probability that fixed-step gradient descent becomes unstable. We further show that the sparsity-induced effect persists under nonlinear activations.
Training-Free Looped Transformers
Chen, Lizhang, Li, Jonathan, Liang, Chen, Lao, Ni, Liu, Qiang
We introduce training-free looped transformers, in which a lightweight inference-time wrapper loops a contiguous mid-stack block of layers of a frozen checkpoint without additional fine-tuning, continued training, or architectural changes. Unlike prior looped transformer methods that train with the looped structure end-to-end, we retrofit recurrence onto pretrained models at test time. We show that naive block reapplication usually degrades performance, highlighting the importance of the loop application strategy. Motivated by viewing a pre-norm transformer block as a forward Euler step on an ODE, we instead treat looping as a refinement of the same approximation, replacing one large update with smaller damped sub-steps. Across seven dense, sparse MoE, and MLA+MoE model families, our method improves Qwen3-4B-Instruct by +2.64 pp on MMLU-Pro, Qwen3-30B-A3B-Instruct by +1.14 pp on CommonsenseQA, and Moonlight-16B-A3B-Instruct by +1.20 pp on OpenBookQA.
Scotland's 'green datacentres' policy ignores emissions impact of AI, analysis shows
Facilities can be branded as aligned with Scotland's climate goals despite significant emissions, said APRS. Facilities can be branded as aligned with Scotland's climate goals despite significant emissions, said APRS. Scotland's'green datacentres' policy ignores emissions impact of AI, analysis shows A Scottish government policy designed to encourage datacentres to build in Scotland could lead to a massive volume of carbon emissions being ignored, according to an analysis by a Scottish charity. "Green datacentres" are at the heart of Scotland's ambitions to develop economically. Enshrined in national policy, they are part of a larger, UK-wide effort to attract big AI investment to Scotland.
How to avoid garbage news on Google Search
'Preferred sources' ensures you're seeing the news outlets you want to see. More information Adding us as a Preferred Source in Google by using this link indicates that you would like to see more of our content in Google News results. Get Google News working the way you want it to. Breakthroughs, discoveries, and DIY tips sent six days a week. When you search Google for something topical, you might see a cluster of headlines from news outlets, reporting breaking stories related to your search query.
Elizabeth Hurley is locked in for summer, hockey goalie Mikayla Demaiter turns up the heat & baseball and meat
Man finds poop on his roof, and if that wasn't bad enough, it led to a mountain lion encounter Sydney Thomas dominates the red carpet in Cannes as her star continues to rise, new MLB power couple & MEAT! Viral staff photo reveals just how bloated Stephen Colbert's'Late Show' operation really was Four of the most controversial television finales in honor of'The Boys' despised ending Sophie Cuningham has heads spinning with her pregame outfit, Colbert's final jab & lessons from Kyle Busch Adrenaline-packed preview released for upcoming D-Day film'Pressure,' features loaded cast Kacey Musgraves responds to'fat activist' furious because she can't fit into her new Walmart clothing line Selena Gomez is reportedly bringing her talents to award-winning director's new four-hour X-rated movie Minka Kelly uncorks a heater at 45, ABS backfires spectacularly and LSU parents vs a security guard! Robot's lifeless corpse hauled off stage after fall during disastrous Michael Jackson impression Bear cubs spar on woman's front porch in adorable viral nature video, reactions pour in Sen Barrasso details Trump's nearly finalized Iran deal, stance on Strait of Hormuz We must'forget our personal differences' and get back to work: Sen Tommy Tuberville They obviously didn't get the memo here about Memorial Day Weekend being unofficial start of summer. It's cool this morning and it's not even supposed to get into the 80s today. But you know who did receive the memo?