Large Language Model
InterFeat: A Pipeline for Finding Interesting Scientific Features
Ofer, Dan, Linial, Michal, Shahaf, Dafna
Finding interesting phenomena is the core of scientific discovery, but it is a manual, ill-defined concept. We present an integrative pipeline for automating the discovery of interesting simple hypotheses (feature-target relations with effect direction and a potential underlying mechanism) in structured biomedical data. The pipeline combines machine learning, knowledge graphs, literature search and Large Language Models. We formalize "interestingness" as a combination of novelty, utility and plausibility. On 8 major diseases from the UK Biobank, our pipeline consistently recovers risk factors years before their appearance in the literature. 40--53% of our top candidates were validated as interesting, compared to 0--7% for a SHAP-based baseline. Overall, 28% of 109 candidates were interesting to medical experts. The pipeline addresses the challenge of operationalizing "interestingness" scalably and for any target. We release data and code: https://github.com/LinialLab/InterFeat
Multi-Modal Multi-Task (M3T) Federated Foundation Models for Embodied AI: Potentials and Challenges for Edge Integration
Borazjani, Kasra, Abdisarabshali, Payam, Nadimi, Fardis, Khosravan, Naji, Liwang, Minghui, Wang, Xianbin, Hong, Yiguang, Hosseinalipour, Seyyedali
As embodied AI systems become increasingly multi-modal, personalized, and interactive, they must learn effectively from diverse sensory inputs, adapt continually to user preferences, and operate safely under resource and privacy constraints. These challenges expose a pressing need for machine learning models capable of swift, context-aware adaptation while balancing model generalization and personalization. Here, two methods emerge as suitable candidates, each offering parts of these capabilities: multi-modal multi-task foundation models (M3T-FMs) provide a pathway toward generalization across tasks and modalities, whereas federated learning (FL) offers the infrastructure for distributed, privacy-preserving model updates and user-level model personalization. However, when used in isolation, each of these approaches falls short of meeting the complex and diverse capability requirements of real-world embodied AI environments. In this vision paper, we introduce multi-modal multi-task federated foundation models (M3T-FFMs) for embodied AI, a new paradigm that unifies the strengths of M3T-FMs with the privacy-preserving distributed training nature of FL, enabling intelligent systems at the wireless edge. We collect critical deployment dimensions of M3T-FFMs in embodied AI ecosystems under a unified framework, which we name "EMBODY": Embodiment heterogeneity, Modality richness and imbalance, Bandwidth and compute constraints, On-device continual learning, Distributed control and autonomy, and Yielding safety, privacy, and personalization. For each, we identify concrete challenges and envision actionable research directions. We also present an evaluation framework for deploying M3T-FFMs in embodied AI systems, along with the associated trade-offs. Finally, we present a prototype implementation of M3T-FFMs and evaluate their energy and latency performance.
Assessing and Mitigating Medical Knowledge Drift and Conflicts in Large Language Models
Wu, Weiyi, Xu, Xinwen, Gao, Chongyang, Diao, Xingjian, Li, Siting, Salas, Lucas A., Gui, Jiang
Large Language Models (LLMs) have great potential in the field of health care, yet they face great challenges in adapting to rapidly evolving medical knowledge. This can lead to outdated or contradictory treatment suggestions. This study investigated how LLMs respond to evolving clinical guidelines, focusing on concept drift and internal inconsistencies. We developed the DriftMedQA benchmark to simulate guideline evolution and assessed the temporal reliability of various LLMs. Our evaluation of seven state-of-the-art models across 4,290 scenarios demonstrated difficulties in rejecting outdated recommendations and frequently endorsing conflicting guidance. Additionally, we explored two mitigation strategies: Retrieval-Augmented Generation and preference fine-tuning via Direct Preference Optimization. While each method improved model performance, their combination led to the most consistent and reliable results. These findings underscore the need to improve LLM robustness to temporal shifts to ensure more dependable applications in clinical practice. The dataset is available at https://huggingface.co/datasets/RDBH/DriftMed.
Test It Before You Trust It: Applying Software Testing for Trustworthy In-context Learning
Racharak, Teeradaj, Ragkhitwetsagul, Chaiyong, Sontesadisai, Chommakorn, Sunetnanta, Thanwadee
In-context learning (ICL) has emerged as a powerful capability of large language models (LLMs), enabling them to perform new tasks based on a few provided examples without explicit fine-tuning. Despite their impressive adaptability, these models remain vulnerable to subtle adversarial perturbations and exhibit unpredictable behavior when faced with linguistic variations. Inspired by software testing principles, we introduce a software testing-inspired framework, called MMT4NL, for evaluating the trustworthiness of in-context learning by utilizing adversarial perturbations and software testing techniques. It includes diverse evaluation aspects of linguistic capabilities for testing the ICL capabilities of LLMs. MMT4NL is built around the idea of crafting metamorphic adversarial examples from a test set in order to quantify and pinpoint bugs in the designed prompts of ICL. Our philosophy is to treat any LLM as software and validate its functionalities just like testing the software. Finally, we demonstrate applications of MMT4NL on the sentiment analysis and question-answering tasks. Our experiments could reveal various linguistic bugs in state-of-the-art LLMs.
In-context Ranking Preference Optimization
Wu, Junda, Surana, Rohan, Xie, Zhouhang, Shen, Yiran, Xia, Yu, Yu, Tong, Rossi, Ryan A., Ammanabrolu, Prithviraj, McAuley, Julian
Recent developments in Direct Preference Optimization (DPO) allow large language models (LLMs) to function as implicit ranking models by maximizing the margin between preferred and non-preferred responses. In practice, user feedback on such lists typically involves identifying a few relevant items in context rather than providing detailed pairwise comparisons for every possible item pair. Moreover, many complex information retrieval tasks, such as conversational agents and summarization systems, critically depend on ranking the highest-quality outputs at the top, emphasizing the need to support natural and flexible forms of user feedback. To address the challenge of limited and sparse pairwise feedback in the in-context setting, we propose an In-context Ranking Preference Optimization (IRPO) framework that directly optimizes LLMs based on ranking lists constructed during inference. To further capture flexible forms of feedback, IRPO extends the DPO objective by incorporating both the relevance of items and their positions in the list. Modeling these aspects jointly is non-trivial, as ranking metrics are inherently discrete and non-differentiable, making direct optimization difficult. To overcome this, IRPO introduces a differentiable objective based on positional aggregation of pairwise item preferences, enabling effective gradient-based optimization of discrete ranking metrics. We further provide theoretical insights showing that IRPO (i) automatically emphasizes items with greater disagreement between the model and the reference ranking, and (ii) links its gradient to an importance sampling estimator, yielding an unbiased estimator with reduced variance. Empirical results show IRPO outperforms standard DPO approaches in ranking performance, highlighting its effectiveness in aligning LLMs with direct in-context ranking preferences.
OpenDeception: Benchmarking and Investigating AI Deceptive Behaviors via Open-ended Interaction Simulation
Wu, Yichen, Pan, Xudong, Hong, Geng, Yang, Min
As the general capabilities of large language models (LLMs) improve and agent applications become more widespread, the underlying deception risks urgently require systematic evaluation and effective oversight. Unlike existing evaluation which uses simulated games or presents limited choices, we introduce OpenDeception, a novel deception evaluation framework with an open-ended scenario dataset. OpenDeception jointly evaluates both the deception intention and capabilities of LLM-based agents by inspecting their internal reasoning process. Specifically, we construct five types of common use cases where LLMs intensively interact with the user, each consisting of ten diverse, concrete scenarios from the real world. To avoid ethical concerns and costs of high-risk deceptive interactions with human testers, we propose to simulate the multi-turn dialogue via agent simulation. Extensive evaluation of eleven mainstream LLMs on OpenDeception highlights the urgent need to address deception risks and security concerns in LLM-based agents: the deception intention ratio across the models exceeds 80%, while the deception success rate surpasses 50%. Furthermore, we observe that LLMs with stronger capabilities do exhibit a higher risk of deception, which calls for more alignment efforts on inhibiting deceptive behaviors.
Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models
NVIDIA, null, :, null, Blakeman, Aaron, Basant, Aarti, Khattar, Abhinav, Renduchintala, Adithya, Bercovich, Akhiad, Ficek, Aleksander, Bjorlin, Alexis, Taghibakhshi, Ali, Deshmukh, Amala Sanjay, Mahabaleshwarkar, Ameya Sunil, Tao, Andrew, Shors, Anna, Aithal, Ashwath, Poojary, Ashwin, Dattagupta, Ayush, Buddharaju, Balaram, Chen, Bobby, Ginsburg, Boris, Wang, Boxin, Norick, Brandon, Butterfield, Brian, Catanzaro, Bryan, del Mundo, Carlo, Dong, Chengyu, Harvey, Christine, Parisien, Christopher, Su, Dan, Korzekwa, Daniel, Yin, Danny, Gitman, Daria, Mosallanezhad, David, Narayanan, Deepak, Fridman, Denys, Rekesh, Dima, Ma, Ding, Pykhtar, Dmytro, Ahn, Dong, Riach, Duncan, Stosic, Dusan, Long, Eileen, Segal, Elad, Evans, Ellie, Chung, Eric, Galinkin, Erick, Bakhturina, Evelina, Dobrowolska, Ewa, Jia, Fei, Liu, Fuxiao, Prasad, Gargi, Shen, Gerald, Liu, Guilin, Chen, Guo, Qian, Haifeng, Ngo, Helen, Liu, Hongbin, Li, Hui, Gitman, Igor, Karmanov, Ilia, Moshkov, Ivan, Golan, Izik, Kautz, Jan, Scowcroft, Jane Polak, Casper, Jared, Seppanen, Jarno, Lu, Jason, Sewall, Jason, Zeng, Jiaqi, You, Jiaxuan, Zhang, Jimmy, Zhang, Jing, Huang, Jining, Xue, Jinze, Huang, Jocelyn, Conway, Joey, Kamalu, John, Barker, Jon, Cohen, Jonathan, Jennings, Joseph, Parmar, Jupinder, Sapra, Karan, Briski, Kari, Chumachenko, Kateryna, Luna, Katherine, Santhanam, Keshav, Kong, Kezhi, Sivamani, Kirthi, Pawelec, Krzysztof, Anik, Kumar, Li, Kunlun, McAfee, Lawrence, Derczynski, Leon, Pavao, Lindsey, Vega, Luis, Voegtle, Lukas, Bala, Maciej, de Melo, Maer Rodrigues, Sreedhar, Makesh Narsimhan, Chochowski, Marcin, Kliegl, Markus, Stepniewska-Dziubinska, Marta, Le, Matthieu, Novikov, Matvei, Samadi, Mehrzad, Andersch, Michael, Evans, Michael, Martinez, Miguel, Chrzanowski, Mike, Ranzinger, Mike, Blaz, Mikolaj, Smelyanskiy, Misha, Fawzy, Mohamed, Shoeybi, Mohammad, Patwary, Mostofa, Lee, Nayeon, Tajbakhsh, Nima, Xu, Ning, Rybakov, Oleg, Kuchaiev, Oleksii, Delalleau, Olivier, Nitski, Osvald, Chadha, Parth, Shamis, Pasha, Micikevicius, Paulius, Molchanov, Pavlo, Dykas, Peter, Fischer, Philipp, Aquilanti, Pierre-Yves, Bialecki, Piotr, Varshney, Prasoon, Gundecha, Pritam, Tredak, Przemek, Karimi, Rabeeh, Kandu, Rahul, El-Yaniv, Ran, Joshi, Raviraj, Waleffe, Roger, Zhang, Ruoxi, Kavanaugh, Sabrina, Jain, Sahil, Kriman, Samuel, Lym, Sangkug, Satheesh, Sanjeev, Muralidharan, Saurav, Narenthiran, Sean, Anandaraj, Selvaraj, Bak, Seonmyeong, Kashirsky, Sergey, Han, Seungju, Acharya, Shantanu, Ghosh, Shaona, Sreenivas, Sharath Turuvekere, Clay, Sharon, Thomas, Shelby, Prabhumoye, Shrimai, Pachori, Shubham, Toshniwal, Shubham, Prayaga, Shyamala, Jain, Siddhartha, Das, Sirshak, Kierat, Slawek, Majumdar, Somshubra, Han, Song, Singhal, Soumye, Niverty, Sriharsha, Alborghetti, Stefania, Panguluri, Suseella, Bhendigeri, Swetha, Akter, Syeda Nahida, Migacz, Szymon, Shiri, Tal, Kong, Terry, Roman, Timo, Ronen, Tomer, Saar, Trisha, Konuk, Tugrul, Rintamaki, Tuomas, Poon, Tyler, De, Ushnish, Noroozi, Vahid, Singh, Varun, Korthikanti, Vijay, Kurin, Vitaly, Ahmad, Wasi Uddin, Du, Wei, Ping, Wei, Dai, Wenliang, Byeon, Wonmin, Ren, Xiaowei, Xu, Yao, Choi, Yejin, Zhang, Yian, Lin, Ying, Suhara, Yoshi, Yu, Zhiding, Li, Zhiqi, Li, Zhiyu, Zhu, Zhongbo, Yang, Zhuolin, Chen, Zijia
As inference-time scaling becomes critical for enhanced reasoning capabilities, it is increasingly becoming important to build models that are efficient to infer. We introduce Nemotron-H, a family of 8B and 56B/47B hybrid Mamba-Transformer models designed to reduce inference cost for a given accuracy level. To achieve this goal, we replace the majority of self-attention layers in the common Transformer model architecture with Mamba layers that perform constant computation and require constant memory per generated token. We show that Nemotron-H models offer either better or on-par accuracy compared to other similarly-sized state-of-the-art open-sourced Transformer models (e.g., Qwen-2.5-7B/72B and Llama-3.1-8B/70B), while being up to 3$\times$ faster at inference. To further increase inference speed and reduce the memory required at inference time, we created Nemotron-H-47B-Base from the 56B model using a new compression via pruning and distillation technique called MiniPuzzle. Nemotron-H-47B-Base achieves similar accuracy to the 56B model, but is 20% faster to infer. In addition, we introduce an FP8-based training recipe and show that it can achieve on par results with BF16-based training. This recipe is used to train the 56B model. We are releasing Nemotron-H base model checkpoints with support in Hugging Face and NeMo.
Hallucination Detection on a Budget: Efficient Bayesian Estimation of Semantic Entropy
Ciosek, Kamil, Felicioni, Nicolรฒ, Ghiassian, Sina
Detecting whether an LLM hallucinates is an important research challenge. One promising way of doing so is to estimate the semantic entropy (Farquhar et al., 2024) of the distribution of generated sequences. We propose a new algorithm for doing that, with two main advantages. First, due to us taking the Bayesian approach, we achieve a much better quality of semantic entropy estimates for a given budget of samples from the LLM. Second, we are able to tune the number of samples adaptively so that `harder' contexts receive more samples. We demonstrate empirically that our approach systematically beats the baselines, requiring only 53% of samples used by Farquhar et al. (2024) to achieve the same quality of hallucination detection as measured by AUROC. Moreover, quite counterintuitively, our estimator is useful even with just one sample from the LLM.
Efficient Dynamic Clustering-Based Document Compression for Retrieval-Augmented-Generation
Li, Weitao, Liu, Kaiming, Zhang, Xiangyu, Lei, Xuanyu, Ma, Weizhi, Liu, Yang
Retrieval-Augmented Generation (RAG) has emerged as a widely adopted approach for knowledge injection during large language model (LLM) inference in recent years. However, due to their limited ability to exploit fine-grained inter-document relationships, current RAG implementations face challenges in effectively addressing the retrieved noise and redundancy content, which may cause error in the generation results. To address these limitations, we propose an Efficient Dynamic Clustering-based document Compression framework (EDC2-RAG) that utilizes latent inter-document relationships while simultaneously removing irrelevant information and redundant content. We validate our approach, built upon GPT-3.5-Turbo and GPT-4o-mini, on widely used knowledge-QA and Hallucination-Detection datasets. Experimental results show that our method achieves consistent performance improvements across various scenarios and experimental settings, demonstrating strong robustness and applicability. Our code and datasets are available at https://github.com/Tsinghua-dhy/EDC-2-RAG.
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding
Li, Ziheng, Sun, Zexu, Zhao, Jinman, Min, Erxue, Zeng, Yongcheng, Wu, Hui, Cai, Hengyi, Wang, Shuaiqiang, Yin, Dawei, Chen, Xu, Deng, Zhi-Hong
Reinforcement learning with verifiable rewards (RL VR) has achieved remarkable success in enhancing the reasoning capabilities of large language models (LLMs). However, existing RL VR methods often suffer from exploration inefficiency due to mismatches between the training data's difficulty and the model's capability. LLMs fail to discover viable reasoning paths when problems are overly difficult, while learning little new capability when problems are too simple. Building on this analysis, we propose SEELE, a novel supervision-aided RL VR framework that dynamically adjusts problem difficulty to stay within the high-efficiency region. Unlike previous hint-based approaches, SEELE deliberately and adaptively adjusts the hint length for each problem to achieve an optimal difficulty. To determine the optimal hint length, SEELE employs a multi-round rollout sampling strategy. In each round, it fits an item response theory model to the accuracy-hint pairs collected in preceding rounds to predict the required hint length for the next round. Experimental results show that SEELE outperforms Group Relative Policy Optimization (GRPO) and Supervised Fine-tuning (SFT) by +11.8 and +10.5 points, respectively, and surpasses the best previous supervision-aided approach by +3.6 points on average across six math reasoning benchmarks.