Deep Learning
Get RICH or Die Scaling: Profitably Trading Inference Compute for Robustness
McDonald, Tavish, Lei, Bo, Fort, Stanislav, Kailkhura, Bhavya, Bartoldson, Brian
Models are susceptible to adversarially out-of-distribution (OOD) data despite large training-compute investments into their robustification. Zaremba et al. (2025) make progress on this problem at test time, showing LLM reasoning improves satisfaction of model specifications designed to thwart attacks, resulting in a correlation between reasoning effort and robustness to jailbreaks. However, this benefit of test compute fades when attackers are given access to gradients or multimodal inputs. We address this gap, clarifying that inference-compute offers benefits even in such cases. Our approach argues that compositional generalization, through which OOD data is understandable via its in-distribution (ID) components, enables adherence to defensive specifications on adversarially OOD inputs. Namely, we posit the Robustness from Inference Compute Hypothesis (RICH): inference-compute defenses profit as the model's training data better reflects the attacked data's components. We empirically support this hypothesis across vision language model and attack types, finding robustness gains from test-time compute if specification following on OOD data is unlocked by compositional generalization. For example, InternVL 3.5 gpt-oss 20B gains little robustness when its test compute is scaled, but such scaling adds significant robustness if we first robustify its vision encoder. This correlation of inference-compute's robustness benefit with base model robustness is the rich-get-richer dynamic of the RICH: attacked data components are more ID for robustified models, aiding compositional generalization to OOD data. Thus, we advise layering train-time and test-time defenses to obtain their synergistic benefit.
Flexible Swarm Learning May Outpace Foundation Models in Essential Tasks
Samadi, Moein E., Schuppert, Andreas
Foundation models have rapidly advanced AI, raising the question of whether their decisions will ultimately surpass human strategies in real-world domains. The exponential, and possibly super-exponential, pace of AI development makes such analysis elusive. Nevertheless, many application areas that matter for daily life and society show only modest gains so far; a prominent case is diagnosing and treating dynamically evolving disease in intensive care. The common challenge is adapting complex systems to dynamic environments. Effective strategies must optimize outcomes in systems composed of strongly interacting functions while avoiding shared side effects; this requires reliable, self-adaptive modeling. These tasks align with building digital twins of highly complex systems whose mechanisms are not fully or quantitatively understood. It is therefore essential to develop methods for self-adapting AI models with minimal data and limited mechanistic knowledge. As this challenge extends beyond medicine, AI should demonstrate clear superiority in these settings before assuming broader decision-making roles. We identify the curse of dimensionality as a fundamental barrier to efficient self-adaptation and argue that monolithic foundation models face conceptual limits in overcoming it. As an alternative, we propose a decentralized architecture of interacting small agent networks (SANs). We focus on agents representing the specialized substructure of the system, where each agent covers only a subset of the full system functions. Drawing on mathematical results on the learning behavior of SANs and evidence from existing applications, we argue that swarm-learning in diverse swarms can enable self-adaptive SANs to deliver superior decision-making in dynamic environments compared with monolithic foundation models, though at the cost of reduced reproducibility in detail.
Auditing Algorithmic Bias in Transformer-Based Trading
Gerami, Armin, Duraiswami, Ramani
Transformer models have become increasingly popular in financial applications, yet their potential risk making and biases remain under-explored. The purpose of this work is to audit the reliance of the model on volatile data for decision-making, and quantify how the frequency of price movements affects the model's prediction confidence. We employ a transformer model for prediction, and introduce a metric based on Partial Information Decomposition (PID) to measure the influence of each asset on the model's decision making. Our analysis reveals two key observations: first, the model disregards data volatility entirely, and second, it is biased toward data with lower-frequency price movements.
SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations
Liang, Buyun, Peng, Liangzu, Luo, Jinqi, Thaker, Darshan, Chan, Kwan Ho Ryan, Vidal, Renรฉ
Large Language Models (LLMs) are increasingly deployed in high-risk domains. However, state-of-the-art LLMs often produce hallucinations, raising serious concerns about their reliability. Prior work has explored adversarial attacks for hallucination elicitation in LLMs, but it often produces unrealistic prompts, either by inserting gibberish tokens or by altering the original meaning. As a result, these approaches offer limited insight into how hallucinations may occur in practice. While adversarial attacks in computer vision often involve realistic modifications to input images, the problem of finding realistic adversarial prompts for eliciting LLM hallucinations has remained largely underexplored. To address this gap, we propose Semantically Equivalent and Coherent Attacks (SECA) to elicit hallucinations via realistic modifications to the prompt that preserve its meaning while maintaining semantic coherence. Our contributions are threefold: (i) we formulate finding realistic attacks for hallucination elicitation as a constrained optimization problem over the input prompt space under semantic equivalence and coherence constraints; (ii) we introduce a constraint-preserving zeroth-order method to effectively search for adversarial yet feasible prompts; and (iii) we demonstrate through experiments on open-ended multiple-choice question answering tasks that SECA achieves higher attack success rates while incurring almost no semantic equivalence or semantic coherence errors compared to existing methods. SECA highlights the sensitivity of both open-source and commercial gradient-inaccessible LLMs to realistic and plausible prompt variations. Code is available at https://github.com/Buyun-Liang/SECA.
Beyond Token Length: Step Pruner for Efficient and Accurate Reasoning in Large Language Models
Wu, Canhui, Cao, Qiong, Li, Chang, Wang, Zhenfang, Xue, Chao, Fan, Yuwei, Xi, Wei, He, Xiaodong
Large Reasoning Models (LRMs) demonstrate strong performance on complex tasks but often suffer from excessive verbosity, known as "overthinking." Existing solutions via reinforcement learning (RL) typically penalize generated tokens to promote conciseness. However, these methods encounter two challenges: responses with fewer tokens do not always correspond to fewer reasoning steps, and models may develop hacking behavior in later stages of training by discarding reasoning steps to minimize token usage. In this work, we introduce \textbf{Step Pruner (SP)}, an RL framework that steers LRMs toward more efficient reasoning by favoring compact reasoning steps. Our step-aware reward function prioritizes correctness while imposing penalties for redundant steps, and withholds rewards for incorrect responses to prevent the reinforcement of erroneous reasoning. Moreover, we propose a dynamic stopping mechanism: when the model's output no longer shortens, training is halted to prevent hacking behavior caused by the merging of steps. Extensive experiments across four reasoning benchmarks demonstrate that SP achieves state-of-the-art accuracy while significantly reducing response length. For instance, on AIME24, SP reduces token usage by \textbf{69.7\%}.
Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer
Gemini Robotics Team, null, Abdolmaleki, Abbas, Abeyruwan, Saminda, Ainslie, Joshua, Alayrac, Jean-Baptiste, Arenas, Montserrat Gonzalez, Balakrishna, Ashwin, Batchelor, Nathan, Bewley, Alex, Bingham, Jeff, Bloesch, Michael, Bousmalis, Konstantinos, Brakel, Philemon, Brohan, Anthony, Buschmann, Thomas, Byravan, Arunkumar, Cabi, Serkan, Caluwaerts, Ken, Casarini, Federico, Chan, Christine, Chang, Oscar, Chappellet-Volpini, London, Chen, Jose Enrique, Chen, Xi, Chiang, Hao-Tien Lewis, Choromanski, Krzysztof, Collister, Adrian, D'Ambrosio, David B., Dasari, Sudeep, Davchev, Todor, Dave, Meet Kirankumar, Devin, Coline, Di Palo, Norman, Ding, Tianli, Doersch, Carl, Dostmohamed, Adil, Du, Yilun, Dwibedi, Debidatta, Egambaram, Sathish Thoppay, Elabd, Michael, Erez, Tom, Fang, Xiaolin, Fantacci, Claudio, Fong, Cody, Frey, Erik, Fu, Chuyuan, Gao, Ruiqi, Giustina, Marissa, Gopalakrishnan, Keerthana, Graesser, Laura, Groth, Oliver, Gupta, Agrim, Hafner, Roland, Hansen, Steven, Hasenclever, Leonard, Haves, Sam, Heess, Nicolas, Hernaez, Brandon, Hofer, Alex, Hsu, Jasmine, Huang, Lu, Huang, Sandy H., Iscen, Atil, Jacob, Mithun George, Jain, Deepali, Jesmonth, Sally, Jindal, Abhishek, Julian, Ryan, Kalashnikov, Dmitry, Karagozler, M. Emre, Karp, Stefani, Kecman, Matija, Kew, J. Chase, Kim, Donnie, Kim, Frank, Kim, Junkyung, Kipf, Thomas, Kirmani, Sean, Konyushkova, Ksenia, Ku, Li Yang, Kuang, Yuheng, Lampe, Thomas, Laurens, Antoine, Le, Tuan Anh, Leal, Isabel, Lee, Alex X., Lee, Tsang-Wei Edward, Lever, Guy, Liang, Jacky, Lin, Li-Heng, Liu, Fangchen, Long, Shangbang, Lu, Caden, Maddineni, Sharath, Majumdar, Anirudha, Maninis, Kevis-Kokitsi, Marmon, Andrew, Martinez, Sergio, Michaely, Assaf Hurwitz, Milonopoulos, Niko, Moore, Joss, Moreno, Robert, Neunert, Michael, Nori, Francesco, Ortiz, Joy, Oslund, Kenneth, Parada, Carolina, Parisotto, Emilio, Paryag, Amaris, Pooley, Acorn, Power, Thomas, Quaglino, Alessio, Qureshi, Haroon, Raju, Rajkumar Vasudeva, Ran, Helen, Rao, Dushyant, Rao, Kanishka, Reid, Isaac, Rendleman, David, Reymann, Krista, Rivas, Miguel, Romano, Francesco, Rubanova, Yulia, Sampedro, Peter Pastor, Sanketi, Pannag R, Shah, Dhruv, Sharma, Mohit, Shea, Kathryn, Shridhar, Mohit, Shu, Charles, Sindhwani, Vikas, Singh, Sumeet, Soricut, Radu, Sterneck, Rachel, Storz, Ian, Surdulescu, Razvan, Tan, Jie, Tompson, Jonathan, Tunyasuvunakool, Saran, Varley, Jake, Vesom, Grace, Vezzani, Giulia, Villalonga, Maria Bauza, Vinyals, Oriol, Wagner, Renรฉ, Wahid, Ayzaan, Welker, Stefan, Wohlhart, Paul, Wu, Chengda, Wulfmeier, Markus, Xia, Fei, Xiao, Ted, Xie, Annie, Xie, Jinyu, Xu, Peng, Xu, Sichun, Xu, Ying, Xu, Zhuo, Yan, Jimmy, Yang, Sherry, Yang, Skye, Yang, Yuxiang, Yu, Hiu Hong, Yu, Wenhao, Yuan, Wentao, Yuan, Yuan, Zhang, Jingwei, Zhang, Tingnan, Zhang, Zhiyuan, Zhou, Allan, Zhou, Guangyao, Zhou, Yuxiang
General-purpose robots need a deep understanding of the physical world, advanced reasoning, and general and dexterous control. This report introduces the latest generation of the Gemini Robotics model family: Gemini Robotics 1.5, a multi-embodiment Vision-Language-Action (VLA) model, and Gemini Robotics-ER 1.5, a state-of-the-art Embodied Reasoning (ER) model. We are bringing together three major innovations. First, Gemini Robotics 1.5 features a novel architecture and a Motion Transfer (MT) mechanism, which enables it to learn from heterogeneous, multi-embodiment robot data and makes the VLA more general. Second, Gemini Robotics 1.5 interleaves actions with a multi-level internal reasoning process in natural language. This enables the robot to "think before acting" and notably improves its ability to decompose and execute complex, multi-step tasks, and also makes the robot's behavior more interpretable to the user. Third, Gemini Robotics-ER 1.5 establishes a new state-of-the-art for embodied reasoning, i.e., for reasoning capabilities that are critical for robots, such as visual and spatial understanding, task planning, and progress estimation. Together, this family of models takes us a step towards an era of physical agents-enabling robots to perceive, think and then act so they can solve complex multi-step tasks.
Learning to Generate Rigid Body Interactions with Video Diffusion Models
Romero, David, Bermudez, Ariana, Li, Hao, Pizzati, Fabio, Laptev, Ivan
Recent video generation models have achieved remarkable progress and are now deployed in film, social media production, and advertising. Beyond their creative potential, such models also hold promise as world simulators for robotics and embodied decision making. Despite strong advances, however, current approaches still struggle to generate physically plausible object interactions and lack object-level control mechanisms. To address these limitations, we introduce KineMask, an approach for video generation that enables realistic rigid body control, interactions, and effects. Given a single image and a specified object velocity, our method generates videos with inferred motions and future object interactions. We propose a two-stage training strategy that gradually removes future motion supervision via object masks. Using this strategy we train video diffusion models (VDMs) on synthetic scenes of simple interactions and demonstrate significant improvements of object interactions in real scenes. Furthermore, KineMask integrates low-level motion control with high-level textual conditioning via predicted scene descriptions, leading to support for synthesis of complex dynamical phenomena. Our experiments show that KineMask achieves strong improvements over recent models of comparable size. Ablation studies further highlight the complementary roles of low- and high-level conditioning in VDMs. Our code, model, and data will be made publicly available. Project Page: https://daromog.github.io/KineMask/
REBot: From RAG to CatRAG with Semantic Enrichment and Graph Routing
Ma, Thanh, La, Tri-Tam, Huu, Lam-Thu Le, Nguyen, Minh-Nghi, Luu, Khanh-Van Pham
Academic regulation advising is essential for helping students interpret and comply with institutional policies, yet building effective systems requires domain specific regulatory resources. To address this challenge, we propose REBot, an LLM enhanced advisory chatbot powered by CatRAG, a hybrid retrieval reasoning framework that integrates retrieval augmented generation with graph based reasoning. CatRAG unifies dense retrieval and graph reasoning, supported by a hierarchical, category labeled knowledge graph enriched with semantic features for domain alignment. A lightweight intent classifier routes queries to the appropriate retrieval modules, ensuring both factual accuracy and contextual depth. We construct a regulation specific dataset and evaluate REBot on classification and question answering tasks, achieving state of the art performance with an F1 score of 98.89%. Finally, we implement a web application that demonstrates the practical value of REBot in real world academic advising scenarios.
Exploring System 1 and 2 communication for latent reasoning in LLMs
Coda-Forno, Julian, Zhao, Zhuokai, Zhang, Qiang, Tamboli, Dipesh, Li, Weiwei, Fan, Xiangjun, Zhang, Lizhu, Schulz, Eric, Tseng, Hsiao-Ping
Should LLM reasoning live in a separate module, or within a single model's forward pass and representational space? We study dual-architecture latent reasoning, where a fluent Base exchanges latent messages with a Coprocessor, and test two hypotheses aimed at improving latent communication over Liu et al. (2024): (H1) increase channel capacity; (H2) learn communication via joint finetuning. Under matched latent-token budgets on GPT-2 and Qwen-3, H2 is consistently strongest while H1 yields modest gains. A unified soft-embedding baseline, a single model with the same forward pass and shared representations, using the same latent-token budget, nearly matches H2 and surpasses H1, suggesting current dual designs mostly add compute rather than qualitatively improving reasoning. Across GSM8K, ProsQA, and a Countdown stress test with increasing branching factor, scaling the latent-token budget beyond small values fails to improve robustness. Latent analyses show overlapping subspaces with limited specialization, consistent with weak reasoning gains. We conclude dual-model latent reasoning remains promising in principle, but likely requires objectives and training schedules that explicitly shape latent spaces for algorithmic planning.
Adaptive Canonicalization with Application to Invariant Anisotropic Geometric Networks
Lin, Ya-Wei Eileen, Levie, Ron
Canonicalization is a widely used strategy in equivariant machine learning, enforcing symmetry in neural networks by mapping each input to a standard form. Yet, it often introduces discontinuities that can affect stability during training, limit generalization, and complicate universal approximation theorems. In this paper, we address this by introducing adaptive canonicalization, a general framework in which the canonicalization depends both on the input and the network. Specifically, we present the adaptive canonicalization based on prior maximization, where the standard form of the input is chosen to maximize the predictive confidence of the network. We prove that this construction yields continuous and symmetry-respecting models that admit universal approximation properties. We propose two applications of our setting: (i) resolving eigenbasis ambiguities in spectral graph neural networks, and (ii) handling rotational symmetries in point clouds. We empirically validate our methods on molecular and protein classification, as well as point cloud classification tasks. Our adaptive canonicalization outperforms the three other common solutions to equivariant machine learning: data augmentation, standard canonicalization, and equivariant architectures.