Goto

Collaborating Authors

 Industry


Sequences of Logits Reveal the Low Rank Structure of Language Models

arXiv.org Machine Learning

A major problem in the study of large language models is to understand their inherent low-dimensional structure. We introduce an approach to study the low-dimensional structure of language models at a model-agnostic level: as sequential probabilistic models. We first empirically demonstrate that a wide range of modern language models exhibit low-rank structure: in particular, matrices built from the model's logits for varying sets of prompts and responses have low approximate rank. We then show that this low-rank structure can be leveraged for generation -- in particular, we can generate a response to a target prompt using a linear combination of the model's outputs on unrelated, or even nonsensical prompts. On the theoretical front, we observe that studying the approximate rank of language models in the sense discussed above yields a simple universal abstraction whose theoretical predictions parallel our experiments. We then analyze the representation power of the abstraction and give provable learning guarantees.


Tree Ensemble Explainability through the Hoeffding Functional Decomposition and TreeHFD Algorithm

arXiv.org Machine Learning

Tree ensembles have demonstrated state-of-the-art predictive performance across a wide range of problems involving tabular data. Nevertheless, the black-box nature of tree ensembles is a strong limitation, especially for applications with critical decisions at stake. The Hoeffding or ANOVA functional decomposition is a powerful explainability method, as it breaks down black-box models into a unique sum of lower-dimensional functions, provided that input variables are independent. In standard learning settings, input variables are often dependent, and the Hoeffding decomposition is generalized through hierarchical orthogonality constraints. Such generalization leads to unique and sparse decompositions with well-defined main effects and interactions. However, the practical estimation of this decomposition from a data sample is still an open problem. Therefore, we introduce the TreeHFD algorithm to estimate the Hoeffding decomposition of a tree ensemble from a data sample. We show the convergence of TreeHFD, along with the main properties of orthogonality, sparsity, and causal variable selection. The high performance of TreeHFD is demonstrated through experiments on both simulated and real data, using our treehfd Python package (https://github.com/ThalesGroup/treehfd). Besides, we empirically show that the widely used TreeSHAP method, based on Shapley values, is strongly connected to the Hoeffding decomposition.


From Linear to Nonlinear: Provable Weak-to-Strong Generalization through Feature Learning

arXiv.org Machine Learning

Weak-to-strong generalization refers to the phenomenon where a stronger model trained under supervision from a weaker one can outperform its teacher. While prior studies aim to explain this effect, most theoretical insights are limited to abstract frameworks or linear/random feature models. In this paper, we provide a formal analysis of weak-to-strong generalization from a linear CNN (weak) to a two-layer ReLU CNN (strong). We consider structured data composed of label-dependent signals of varying difficulty and label-independent noise, and analyze gradient descent dynamics when the strong model is trained on data labeled by the pretrained weak model. Our analysis identifies two regimes -- data-scarce and data-abundant -- based on the signal-to-noise characteristics of the dataset, and reveals distinct mechanisms of weak-to-strong generalization. In the data-scarce regime, generalization occurs via benign overfitting or fails via harmful overfitting, depending on the amount of data, and we characterize the transition boundary. In the data-abundant regime, generalization emerges in the early phase through label correction, but we observe that overtraining can subsequently degrade performance.


The Sign Estimator: LLM Alignment in the Face of Choice Heterogeneity

arXiv.org Machine Learning

Traditional LLM alignment methods are vulnerable to heterogeneity in human preferences. Fitting a naïve probabilistic model to pairwise comparison data (say over prompt-completion pairs) yields an inconsistent estimate of the population-average utility -a canonical measure of social welfare. We propose a new method, dubbed the sign estimator, that provides a simple, provably consistent, and efficient estimator by replacing cross-entropy with binary classification loss in the aggregation step. This simple modification recovers consistent ordinal alignment under mild assumptions and achieves the first polynomial finite-sample error bounds in this setting. In realistic simulations of LLM alignment using digital twins, the sign estimator substantially reduces preference distortion over a panel of simulated personas, cutting (angular) estimation error by nearly 35% and decreasing disagreement with true population preferences from 12% to 8% compared to standard RLHF. Our method also compares favorably to panel data heuristics that explicitly model user heterogeneity and require tracking individual-level preference data-all while maintaining the implementation simplicity of existing LLM alignment pipelines.


Differential Privacy as a Perk: Federated Learning over Multiple-Access Fading Channels with a Multi-Antenna Base Station

arXiv.org Machine Learning

Federated Learning (FL) is a distributed learning paradigm that preserves privacy by eliminating the need to exchange raw data during training. In its prototypical edge instantiation with underlying wireless transmissions enabled by analog over-the-air computing (AirComp), referred to as \emph{over-the-air FL (AirFL)}, the inherent channel noise plays a unique role of \emph{frenemy} in the sense that it degrades training due to noisy global aggregation while providing a natural source of randomness for privacy-preserving mechanisms, formally quantified by \emph{differential privacy (DP)}. It remains, nevertheless, challenging to effectively harness such channel impairments, as prior arts, under assumptions of either simple channel models or restricted types of loss functions, mostly considering (local) DP enhancement with a single-round or non-convergent bound on privacy loss. In this paper, we study AirFL over multiple-access fading channels with a multi-antenna base station (BS) subject to user-level DP requirements. Despite a recent study, which claimed in similar settings that artificial noise (AN) must be injected to ensure DP in general, we demonstrate, on the contrary, that DP can be gained as a \emph{perk} even \emph{without} employing any AN. Specifically, we derive a novel bound on DP that converges under general bounded-domain assumptions on model parameters, along with a convergence bound with general smooth and non-convex loss functions. Next, we optimize over receive beamforming and power allocations to characterize the optimal convergence-privacy trade-offs, which also reveal explicit conditions in which DP is achievable without compromising training. Finally, our theoretical findings are validated by extensive numerical results.


CANDI: Hybrid Discrete-Continuous Diffusion Models

arXiv.org Machine Learning

While continuous diffusion has shown remarkable success in continuous domains such as image generation, its direct application to discrete data has underperformed compared to purely discrete formulations. This gap is counterintuitive, given that continuous diffusion learns score functions that enable joint evolution across multiple positions. To understand this gap, we introduce token identifiability as an analytical framework for understanding how Gaussian noise corrupts discrete data through two mechanisms: discrete identity corruption and continuous rank degradation. We reveal that these mechanisms scale differently with vocabulary size, creating a temporal dissonance: at noise levels where discrete corruption preserves enough structure for conditional learning, continuous denoising is trivial; at noise levels where continuous denoising is meaningful, discrete corruption destroys nearly all conditional structure. To solve this, we propose CANDI (Continuous ANd DIscrete diffusion), a hybrid framework that decouples discrete and continuous corruption, enabling simultaneous learning of both conditional structure and continuous geometry. We empirically validate the temporal dissonance phenomenon and demonstrate that CANDI successfully avoids it. This unlocks the benefits of continuous diffusion for discrete spaces: on controlled generation, CANDI enables classifier-based guidance with off-the-shelf classifiers through simple gradient addition; on text generation, CANDI outperforms masked diffusion at low NFE, demonstrating the value of learning continuous gradients for discrete spaces. We include the code on the project page available here: https://patrickpynadath1.github.io/candi-lander


Microsoft reports strong earnings as Azure hit by major outage

The Guardian

Microsoft's CEO, Satya Nadella, speaks at the company's annual developer conference in Seattle, Washington. Microsoft's CEO, Satya Nadella, speaks at the company's annual developer conference in Seattle, Washington. Tech giant reports earnings of $3.72 per share day after deal with OpenAI pushed value of company to more than $4tn Microsoft blew off concerns of overspending on AI on Wednesday, reporting elevated earnings even as it faced an outage of its cloud computing service, Azure, and its office software suite, 365. The strong earnings report comes a day after a deal with OpenAI pushed the value of the tech giant to more than $4tn. After its Xbox and investor relations pages went down, the company issued a statement that said: "We are working to address an issue affecting Azure Front Door that is impacting the availability of some services."


Nvidia becomes first 5 trillion firm as AI rally picks up steam

The Japan Times

Nvidia CEO Jensen Huang speaks during an event in Washington on Tuesday. Nvidia achieved a historic $5 trillion market capitalization on Wednesday as CEO Jensen Huang's spree of deals catapults the artificial intelligence frenzy to new heights. The shares closed 3.1% higher at $207.16, propelling Nvidia just over the milestone. It's only been four months since the company cracked the $4 trillion barrier, and the rally has accelerated as Huang forges new agreements to supply companies from Nokia Oyj to Samsung Electronics and Hyundai Motor Group with chips. Nvidia has become the most-important stock in a bull market that's been driven by optimism for AI to revolutionize the global economy.


Trump-Xi meeting: What's at stake and who has the upper hand?

Al Jazeera

Is the US eyeing its next Latin American target? Why is Trump tearing down parts of the White House? Trump-Xi meeting: What's at stake and who has the upper hand? United States President Donald Trump expects "a lot of problems" will be solved between Washington and Beijing when he meets China's President Xi Jinping in South Korea for a high-stakes meeting on Thursday, amid growing trade tensions between the two. Relations between the two world powers have been strained in recent years, with Washington and Beijing imposing tit-for-tat trade tariffs topping 100 percent against each other this year, the US restricting its exports of semiconductors vital for artificial intelligence (AI) development and Beijing restricting exports of critical rare-earth metals which are vital for the defence industry and also the development of AI, among other issues. On the sidelines of the Asia-Pacific Economic Cooperation (APEC) summit in Gyeongju, South Korea, on Wednesday, Trump said an expected trade deal between China and the US would be good for both countries and "something very exciting for everybody".


As Trump Weighs Sale of Advanced A.I. Chips to China, Critics Sound Alarm

NYT > Economy

As President Trump flew to South Korea on Wednesday to prepare for a summit with the Chinese leader, Xi Jinping, he made some remarks that set off alarm bells among Washington officials concerned about America's rivalry with China. "We'll be speaking about Blackwell," Mr. Trump said of his meeting with Mr. Xi, referring to the most advanced artificial intelligence chip from the U.S. chipmaker Nvidia. Mr. Trump called the technology a "super duper chip"; complimented Nvidia's chief executive, Jensen Huang; and declared, "We're about 10 years ahead of anybody else in chips." Mr. Trump's comments signaled a major potential change for U.S. policy that many Washington officials warn poses a national security risk. Selling such advanced A.I. chips to China is currently banned, and U.S. officials have worked for years to restrain Beijing's access to the cutting-edge technology.