Goto

Collaborating Authors

 Europe


How Prince George will follow in his father's footsteps at Eton College

BBC News

How Prince George will follow in his father's footsteps at Eton College Prince George will attend Eton College in Berkshire from September, Kensington Palace has announced. His father, Prince William, also attended the elite boarding school for boys, where fees are about £63,000 a year. He is the Prince and Princess of Wales' oldest child and the second in line of succession to the throne. The BBC's senior royal correspondent Daniela Relph explains the Royal Family's connection to Eton College. Did Coppell lose his keys during 106 celebration?


France to ditch Palantir's AI data tools in favour of domestic provider

The Guardian

The French decision to use its own AI models comes amid growing concern among European governments about US-controlled technology. The French decision to use its own AI models comes amid growing concern among European governments about US-controlled technology. Move to ChapsVision is to avoid'strategic dependencies', says PM amid concern about reliance on US-controlled tools Tue 16 Jun 2026 13.08 EDTLast modified on Tue 16 Jun 2026 15.39 EDT France's domestic intelligence service is to ditch AI data tools from the US tech company Palantir in favour of a domestic provider in an effort to avoid "strategic dependency", the prime minister, Sébastien Lecornu, has said. "We must use our own AI models; we cannot accept new strategic dependencies in the digital sphere," Lecornu posted on social media. "We cannot rely on tools developed by foreign powers. France must have its own tools."


GRASS: Scalable Data Attribution with Gradient Sparsification and Sparse Projection

Neural Information Processing Systems

Gradient-based data attribution methods, such as influence functions, are critical for understanding the impact of individual training samples without requiring repeated model retraining. However, their scalability is often limited by the high computational and memory costs associated with per-sample gradient computation. In this work, we propose GRASS, a novel gradient compression algorithm and its variants FACTGRASS for linear layers specifically, that explicitly leverage the inherent sparsity of per-sample gradients to achieve sub-linear space and time complexity. Extensive experiments demonstrate the effectiveness of our approach, achieving substantial speedups while preserving data influence fidelity. In particular, FACTGRASS achieves up to 165% faster throughput on billion-scale models compared to the previous state-of-the-art baselines.


What in Common Models Hallucinate When Reasoning Across Scenes

Neural Information Processing Systems

Multimodal language models possess a remarkable ability to handle an openvocabulary worth of objects. Yet the best models still suffer from hallucinations when reasoning about scenes in the real world, revealing a gap between their seemingly strong performance on existing perception benchmarks that are saturating and their reasoning in the real world. To address this gap, we build a novel benchmark of in-the-wild scenes that we call Common-OBench. With more than 10.5k examples using exclusively new images not found in web training data to avoid contamination, Common-OBenchgoes beyond just perception, inspired by cognitive tests for humans, to probe reasoning across scenes by asking "what's in common?". We evaluate leading multimodal language models, including models specifically trained to reason. We find that perceiving objects in single images is easy for most models, yet reasoning across scenes is very challenging even for the best models, including reasoning models. Despite saturating many leaderboards focusing on perception, the best performing model only achieves 35% on Common-OBench--and on Common-OComplex, consisting of more complex scenes, the best model achieves only 1%. Curiously, we find models are more prone to hallucinate when similar objects are present in the scene, suggesting models may be relying on object co-occurrence seen during training. Among the models we evaluated, we found scale can provide modest improvements while models explicitly trained with multi-image inputs show bigger improvements, suggesting scaled multi-image training may offer promise.


Uni-LoRA: One Vector is All You Need

Neural Information Processing Systems

Low-Rank Adaptation (LoRA) has become the de facto parameter-efficient finetuning (PEFT) method for large language models (LLMs) by constraining weight updates to low-rank matrices. Recent works such as Tied-LoRA, VeRA, and VBLoRA push efficiency further by introducing additional constraints to reduce the trainable parameter space. In this paper, we show that the parameter space reduction strategies employed by these LoRA variants can be formulated within a unified framework, Uni-LoRA, where the LoRA parameter space, flattened as a highdimensional vector space RD, can be reconstructed through a projection from a subspace Rd, with d D. We demonstrate that the fundamental difference among various LoRA methods lies in the choice of the projection matrix, P RD d. Most existing LoRA variants rely on layer-wise or structure-specific projections that limit cross-layer parameter sharing, thereby compromising parameter efficiency. In light of this, we introduce an efficient and theoretically grounded projection matrix that is isometric, enabling global parameter sharing and reducing computation overhead. Furthermore, under the unified view of Uni-LoRA, this design requires only a single trainable vector to reconstruct LoRA parameters for the entire LLM - making UniLoRA both a unified framework and a "one-vector-only" solution. Extensive experiments on GLUE, mathematical reasoning, and instruction tuning benchmarks demonstrate that Uni-LoRA achieves state-of-the-art parameter efficiency while outperforming or matching prior approaches in predictive performance.


Russian warship fires warning shots near UK-registered yacht in Channel

BBC News

'It was surreal': British couple describe having warning shots fired near them by Russian warship A retired British couple who were on a yacht which had warning shots fired near it by a Russian warship in the English Channel have told the BBC the experience was surreal. Jane and Alan Kelvey were sailing 23 miles (37km) off the Isle of Wight in international waters when they came into close contact with the Russian frigate, the Admiral Grigorovich on Tuesday. Sir Keir Starmer said firing shots into the path of a UK-registered yacht was reckless - an incident the Ministry of Defence has described as an isolated one. Russia's Defence Ministry said the yacht had been on a dangerous approach towards the warship but the couple said they were not on a collision course. The incident comes days after Royal Marine Commandos intercepted a Russian shadow fleet tanker carrying sanctioned oil in the Channel on Sunday, in the first operation of its kind carried out by the British military.



HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages

Neural Information Processing Systems

Preference datasets are essential for training general-domain, instruction-following language models with Reinforcement Learning from Human Feedback (RLHF). Each subsequent data release raises expectations for future data collection, meaning there is a constant need to advance the quality and diversity of openly available preference data. To address this need, we introduce HelpSteer3-Preference, a permissively licensed (CC-BY-4.0),


Accelerating Chain of Thought Reasoning through Semantically Aligned Implicit Tokens

Neural Information Processing Systems

Chain-of-Thought (CoT) enhances the performance of Large Language Models (LLMs) on reasoning tasks by encouraging step-by-step solutions. However, the verbosity of CoT reasoning hinders its mass deployment in efficiency-critical applications. Recently, implicit CoT approaches have emerged, which encode reasoning steps within LLM's hidden embeddings (termed "implicit reasoning") rather than explicit tokens. This approach accelerates CoT reasoning by reducing the reasoning length and bypassing some LLM components. However, existing implicit CoT methods face two significant challenges: (1) they fail to preserve the semantic alignment between the implicit reasoning (when transformed to natural language) and the ground-truth reasoning, resulting in a significant CoT performance degradation, and (2) they focus on reducing the length of the implicit reasoning; however, they neglect the considerable time cost for an LLM to generate one individual implicit reasoning token.


The Structure of Relation Decoding Linear Operators in Large Language Models

Neural Information Processing Systems

This paper investigates the structure of linear operators introduced in Hernandez et al. [2023] that decode specific relational facts in transformer language models. We extend their single-relation findings to a collection of relations and systematically chart their organization. We show that such collections of relation decoders can be highly compressed by simple order-3 tensor networks without significant loss in decoding accuracy. To explain this surprising redundancy, we develop a cross-evaluation protocol, in which we apply each linear decoder operator to the subjects of every other relation. Our results reveal that these linear maps do not encode distinct relations, but extract recurring, coarse-grained semantic properties (e.g., country of capital city and country of food are both in the country-of-X property). This property-centric structure clarifies both the operators' compressibility and highlights why they generalize only to new relations that are semantically close. Our findings thus interpret linear relational decoding in transformer language models as primarily property-based, rather than relation-specific.1