Goto

Collaborating Authors

 Vienna


Watch the ICRA keynote and plenary talks

Robohub

The 2026 IEEE International Conference on Robotics & Automation (ICRA) was held in Vienna from the 1 - 5 June. The event brought together researchers and industry professionals with expertise spanning many different aspects of robotics. IEEE have made the keynote and plenary talk recordings available to watch . You can also catch the panel discussions and keynote tutorials. Panel 4 - Publish or Perish: Surviving the Paper Deluge - Is AI the solution?


Kamala Harris viciously mocked with resurfaced bumbling 'two words' explanation of AI after desperate call

FOX News

Kamala Harris is being mocked after a resurfaced clip of her rambling AI explanation followed her post calling on Donald Trump to pursue a treaty with China on artificial intelligence.


Europe must build own AI or risk getting cut off by US or China, says ECB's Lagarde

The Guardian

Christine Lagarde spoke in Vienna about Europe and AI on Monday. Christine Lagarde spoke in Vienna about Europe and AI on Monday. Europe must build own AI or risk getting cut off by US or China, says ECB's Lagarde Central bank chief says continent's AI dependency could give trade partners unprecedented leverage in negotiations Europe must develop its own AI technology and build more datacentres in order to nullify the threat of being cut off by the US or China, according to the president of the European Central Bank . Christine Lagarde said the continent needed AI models - the technology that powers AI tools such as chatbots - that were "good enough" to carry out most tasks and run from domestic datacentres. If Europe invests in its own AI tech, said Lagarde, "the threat of being cut off loses its force".



Reflections from ICRA 2026

Robohub

From the 1st-5th June, the robots descended on Vienna. The 2026 IEEE International Conference on Robotics & Automation (ICRA) brought together the top minds in robotics for one short week to showcase the latest technologies, form new collaborations, and exchange ideas. Held at the Messe Wien, a stone's throw from the bank of the Danube, ICRA proved to be equal parts technological marvel and thought-provoking discussion. The host venue for ICRA 2026: Messe Wien, also known as VIECON. My week at ICRA began with the 2nd ICRA 2026 Workshop on Robot Ethics: Ethical, Legal and User Perspectives in Robotics & Automation (WOROBET) .


d7a2222b8d41014e060cfeb0995501d0-Paper-Conference.pdf

Neural Information Processing Systems

How can we trust the correctness of a learned model on a particular input of interest? Model accuracy is typically measured on average over a distribution of inputs, giving no guarantee for any fixed input. This paper proposes a theoreticallyfounded solution to this problem: to train Self-Proving models that prove the correctness of their output to a verification algorithm V via an Interactive Proof. SelfProving models satisfy that, with high probability over an input sampled from a given distribution, the model generates a correct output and successfully proves its correctness to V. The soundness property of V guarantees that, for every input, no model can convince V of the correctness of an incorrect output. Thus, a Self-Proving model proves correctness of most of its outputs, while all incorrect outputs (of any model) are detected by V. We devise and analyze two generic methods for learning Self-Proving models: Transcript Learning (TL) which relies on access to transcripts of accepting interactions, and Reinforcement Learning from Verifier Feedback (RLVF) which trains a model by emulating interactions with the verifier.


AttentionPredictor: Temporal Patterns Matter for KVCache Compression

Neural Information Processing Systems

With the development of large language models (LLMs), efficient inference through Key-Value (KV) cache compression has attracted considerable attention, especially for long-context generation. To compress the KV cache, recent methods identify critical KV tokens through static modeling of attention scores. However, these methods often struggle to accurately determine critical tokens as they neglect the temporal patterns in attention scores, resulting in a noticeable degradation in LLM performance. To address this challenge, we propose AttentionPredictor, which is the first learning-based method to directly predict attention patterns for KV cache compression and critical token identification. Specifically, AttentionPredictor learns a lightweight, unified convolution model to dynamically capture spatiotemporal patterns and predict the next-token attention scores. An appealing feature of AttentionPredictor is that it accurately predicts the attention score and shares the unified prediction model, which consumes negligible memory, among all transformer layers. Moreover, we propose a cross-token critical cache prefetching framework that hides the token estimation time overhead to accelerate the decoding stage. By retaining most of the attention information, AttentionPredictor achieves 13 KV cache compression and 5.6 speedup in a cache offloading scenario with comparable LLM performance, significantly outperforming the stateof-the-arts.


Lua-LLM: Learning Unstructured-Sparsity Allocation for Large Language Models

Neural Information Processing Systems

Large Language Models (LLMs) have demonstrated remarkable capabilities, yet their extensive parameter scales pose significant challenges for practical deployment. Unstructured pruning has emerged as an effective model compression strategy with minimal performance loss, which introduces fine-grained sparsity for weight parameters. While existing methods employ a layer-wise pruning strategy to avoid the complexity of global pruning for billion-scale LLMs, they require appropriate sparsity allocation for the layer-wise pruning objectives and often lead to suboptimal solutions for the overall model. In this paper, we propose Lua-LLM (Learning unstructured-sparsity allocation in LLMs), a learning-based global pruning framework that explores the optimal unstructured sparsity allocation. Unlike existing pruning methods, which primarily focus on allocating per-layer sparsity, Lua-LLM achieves flexible allocation for both layer-wise and intra-layer sparsity.


Domain-Specific Pruning of Large Mixture-of-Experts Models with Few-shot Demonstrations

Neural Information Processing Systems

Mixture-of-Experts (MoE) models achieve a favorable trade-off between performance and inference efficiency by activating only a subset of experts. However, the memory overhead of storing all experts remains a major limitation, especially in large-scale MoE models such as DeepSeek-R1 (671B). In this study, we investigate domain specialization and expert redundancy in large-scale MoE models and uncover a consistent behavior we term few-shot expert localization, with only a few in-domain demonstrations, the model consistently activates a sparse and stable subset of experts on tasks within the same domain. Building on this observation, we propose a simple yet effective pruning framework, EASY-EP, that leverages a few domain-specific demonstrations to identify and retain only the most relevant experts. EASY-EP comprises two key components: output-aware expert importance assessment and expert-level token contribution estimation. The former evaluates the importance of each expert for the current token by considering the gating scores and L2 norm of the outputs of activated experts, while the latter assesses the contribution of tokens based on representation similarities before and after routed experts. Experiments on DeepSeek-R1 and DeepSeek-V3-0324 show that our method can achieve comparable performances and 2.99 throughput under the same memory budget as the full model, with only half the experts.


Factor Decorrelation Enhanced Data Removal from Deep Predictive Models

Neural Information Processing Systems

The imperative of user privacy protection and regulatory compliance necessitates sensitive data removal in model training, yet this process often induces distributional shifts that undermine model performance-particularly in out-of-distribution (OOD) scenarios. To address this issue we propose a novel data removal approach that enhances deep predictive models through factor decorrelation and loss perturbation. Our approach introduces: (1) a discriminative-preserving factor decorrelation module employing dynamic adaptive weight adjustment and iterative representation updating to reduce feature redundancy and minimize inter-feature correlations.