Goto

Collaborating Authors

 Deep Learning


Polar Sparsity: High Throughput Batched LLM Inferencing with Scalable Contextual Sparsity

arXiv.org Artificial Intelligence

Accelerating large language model (LLM) inference is critical for real-world deployments requiring high throughput and low latency. Contextual sparsity, where each token dynamically activates only a small subset of the model parameters, shows promise but does not scale to large batch sizes due to union of active neurons quickly approaching dense computation. We introduce Polar Sparsity, highlighting a key shift in sparsity importance from MLP to Attention layers as we scale batch size and sequence length. While MLP layers become more compute-efficient under batching, their sparsity vanishes. In contrast, attention becomes increasingly more expensive at scale, while their head sparsity remains stable and batch-invariant. We develop Selective Head Attention with hardware-efficient, sparsity-aware GPU kernels, delivering up to \(2.2\times\) end-to-end speedups for models like OPT, LLaMA-2 \& 3, Qwen, Mistral across various batch sizes and sequence lengths without compromising accuracy. To our knowledge, this is the first work to demonstrate that contextual sparsity can scale effectively to large batch sizes, delivering substantial inference acceleration with minimal changes, making Polar Sparsity practical for large-scale, high-throughput LLM deployment systems. Our code is available at: https://github.com/susavlsh10/Polar-Sparsity.


RefiDiff: Progressive Refinement Diffusion for Efficient Missing Data Imputation

arXiv.org Artificial Intelligence

Missing values in high-dimensional, mixed-type datasets pose significant challenges for data imputation, particularly under Missing Not At Random (MNAR) mechanisms. Existing methods struggle to integrate local and global data characteristics, limiting performance in MNAR and high-dimensional settings. We propose an innovative framework, RefiDiff, combining local machine learning predictions with a novel Mamba-based denoising network efficiently capturing long-range dependencies among features and samples with low computational complexity. RefiDiff bridges the predictive and generative paradigms of imputation, leveraging pre-refinement for initial warm-up imputations and post-refinement to polish results, enhancing stability and accuracy. By encoding mixed-type data into unified tokens, RefiDiff enables robust imputation without architectural or hyperparameter tuning. RefiDiff outperforms state-of-the-art (SOT A) methods across missing-value settings, demonstrating strong performance in MNAR settings and superior out-of-sample generalization. Extensive evaluations on nine real-world datasets demonstrate its robustness, scalability, and effectiveness in handling complex missingness patterns.


LExT: Towards Evaluating Trustworthiness of Natural Language Explanations

arXiv.org Artificial Intelligence

As Large Language Models (LLMs) become increasingly integrated into high-stakes domains, there have been several approaches proposed toward generating natural language explanations. These explanations are crucial for enhancing the interpretability of a model, especially in sensitive domains like healthcare, where transparency and reliability are key. In light of such explanations being generated by LLMs and its known concerns, there is a growing need for robust evaluation frameworks to assess model-generated explanations. Natural Language Generation metrics like BLEU and ROUGE capture syntactic and semantic accuracies but overlook other crucial aspects such as factual accuracy, consistency, and faithfulness. To address this gap, we propose a general framework for quantifying trustworthiness of natural language explanations, balancing Plausibility and Faithfulness, to derive a comprehensive Language Explanation Trustworthiness Score (LExT) (The code and set up to reproduce our experiments are publicly available at https://github.com/cerai-iitm/LExT). Applying our domain-agnostic framework to the healthcare domain using public medical datasets, we evaluate six models, including domain-specific and general-purpose models. Our findings demonstrate significant differences in their ability to generate trustworthy explanations. On comparing these explanations, we make interesting observations such as inconsistencies in Faithfulness demonstrated by general-purpose models and their tendency to outperform domain-specific fine-tuned models. This work further highlights the importance of using a tailored evaluation framework to assess natural language explanations in sensitive fields, providing a foundation for improving the trustworthiness and transparency of language models in healthcare and beyond.


Branching Flows: Discrete, Continuous, and Manifold Flow Matching with Splits and Deletions

arXiv.org Machine Learning

Diffusion and flow matching approaches to generative modeling have shown promise in domains where the state space is continuous, such as image generation or protein folding & design, and discrete, exemplified by diffusion large language models. They offer a natural fit when the number of elements in a state is fixed in advance (e.g. images), but require ad hoc solutions when, for example, the length of a response from a large language model, or the number of amino acids in a protein chain is not known a priori. Here we propose Branching Flows, a generative modeling framework that, like diffusion and flow matching approaches, transports a simple distribution to the data distribution. But in Branching Flows, the elements in the state evolve over a forest of binary trees, branching and dying stochastically with rates that are learned by the model. This allows the model to control, during generation, the number of elements in the sequence. We also show that Branching Flows can compose with any flow matching base process on discrete sets, continuous Euclidean spaces, smooth manifolds, and `multimodal' product spaces that mix these components. We demonstrate this in three domains: small molecule generation (multimodal), antibody sequence generation (discrete), and protein backbone generation (multimodal), and show that Branching Flows is a capable distribution learner with a stable learning objective, and that it enables new capabilities.



A beginner's guide to ChatGPT: Make AI work for you

PCWorld

When you purchase through links in our articles, we may earn a small commission. We'll show you how to get started, what you can do, and how to make ChatGPT work for you instead of the other way round. Hardly anyone can have missed the AI phenomenon that has taken the world by storm. Almost every major company has some kind of AI initiative now. Politicians talk about how important it is not to "fall behind in the AI race," and hundreds of millions have started using AI chatbots. The AI wave took off when OpenAI released its chatbot ChatGPT, which gives large language models a conversational interface.


How to Talk to ChatGPT for Free Inside WhatsApp (While You Still Can)

WIRED

Meta's messaging app offers free access to the AI chatbot, but only until January 2026. There are plenty of places you can get access to ChatGPT: Not just in the official apps for the web and mobile devices, but also through Copilot from Microsoft, and in Apple's Siri assistant ... and inside the messaging app WhatsApp . WhatsApp, run by Facebook developer Meta, is available free of charge on the web, and on Android and iOS . It's used by billions of people worldwide, which helps to explain why OpenAI has made ChatGPT available here as well as everywhere else. Unfortunately, OpenAI will be pulling free access to its chatbot within WhatsApp on January 15, 2026.


Improving VMware migration workflows with agentic AI

MIT Technology Review

As licensing costs surge and cloud use becomes more strategic, AI agents are turning months of manual migration work for IT teams into weeks of machine-assisted automation. For years, many chief information officers (CIOs) looked at VMware-to-cloud migrations with a wary pragmatism. Manually mapping dependencies and rewriting legacy apps mid-flight was not an enticing, low-lift proposition for enterprise IT teams. But the calculus for such decisions has changed dramatically in a short period of time. Following recent VMware licensing changes, organizations are seeing greater uncertainty around the platform's future. At the same time, cloud-native innovation is accelerating.


German court rules against OpenAI in copyright case

The Japan Times

The Munich court found that OpenAI, the maker of ChatGPT, was not entitled to use song lyrics to train its artificial intelligence without licenses, and that the artists who wrote them are entitled to compensation. The Munich court found that the maker of ChatGPT was not entitled to use song lyrics to train its artificial intelligence without licenses, and that the artists who wrote them are entitled to compensation. In a time of both misinformation and too much information, quality journalism is more crucial than ever. By subscribing, you can help us get the story right. With your current subscription plan you can comment on stories.


SoftBank sells Nvidia stake for 5.8 billion to fund AI bets

The Japan Times

SoftBank sells Nvidia stake for $5.8 billion to fund AI bets SoftBank Group founder Masayoshi Son is aggressively seeking to capitalize on booming investment in AI and chips, even as he scales back other investments. SoftBank Group sold its entire stake in Nvidia for $5.83 billion to help bankroll artificial intelligence investments, even as investors question the amount of capital pouring into a technology with uncertain returns. Founder Masayoshi Son has been unwinding positions to pay for a plethora of AI projects, from Stargate data centers with OpenAI and Oracle to robot manufacturing sites in the United States. The Nvidia exit coincides with a growing debate about whether spending by big tech firms like Meta Platforms and Alphabet -- expected to surpass $1 trillion in coming years -- will produce commensurate returns. SoftBank's stock slid more than 10% in Tokyo on Wednesday, highlighting how investors remain nervous about lofty tech valuations.