Goto

Collaborating Authors

 Asia


FlexAC: Towards Flexible Control of Associative Reasoning in Multimodal Large Language Models

Neural Information Processing Systems

Multimodal large language models (MLLMs) face an inherent trade-off between faithfulness and creativity, as different tasks require varying degrees of associative reasoning. However, existing methods lack the flexibility to modulate this reasoning strength, limiting MLLMs' adaptability across factual and creative scenarios. To bridge this gap, we propose equipping MLLMs with mechanisms that enable flexible control over associative reasoning. We begin by investigating the internal mechanisms underlying associative behavior in MLLMs and find that: (1) middle layers play a pivotal role in shaping model's associative tendencies, (2) modifying representations in these layers effectively regulates associative reasoning strength, and (3) hallucinations can be exploited to derive steering vectors that guide this modulation. Building on these findings, we introduce Flexible Association Control (FlexAC), a lightweight and training-free framework for modulating associative behavior in MLLMs.


Future Link Prediction Without Memory or Aggregation

Neural Information Processing Systems

Future link prediction on temporal graphs is a fundamental task with wide applicability in real-world dynamic systems. These scenarios often involve both recurring (seen) and novel (unseen) interactions, requiring models to generalize effectively across both types of edges. However, existing methods typically rely on complex memory and aggregation modules, yet struggle to handle unseen edges. In this paper, we revisit the architecture of existing temporal graph models and identify two essential but overlooked modeling requirements for future link prediction: representing nodes with unique identifiers and performing target-aware matching between source and destination nodes. To this end, we propose Cross-Attention based Future Link Predictor on Temporal Graphs (CRAFT), a simple yet effective architecture that discards memory and aggregation modules and instead builds on two components: learnable node embeddings and cross-attention between the destination and the source's recent interactions. This design provides strong expressive power and enables target-aware modeling of the compatibility between candidate destinations and the source's interaction patterns. Extensive experiments on diverse datasets demonstrate that CRAFT consistently achieves superior performance with high efficiency, making it well-suited for large-scale real-world applications.



Spectral Analysis of Diffusion Models with Application to Schedule Design

Neural Information Processing Systems

Diffusion models (DMs) have emerged as powerful tools for modeling complex data distributions and generating realistic new samples. Over the years, advanced architectures and sampling methods have been developed to make these models practically usable. However, certain synthesis process decisions still rely on heuristics without a solid theoretical foundation. In our work, we offer a novel analysis of the DM's inference process, introducing a comprehensive frequency response perspective. Specifically, by relying on Gaussianity assumption, we present the inference process as a closed-form spectral transfer function, capturing how the generated signal evolves in response to the initial noise. We demonstrate how the proposed analysis can be leveraged to design a noise schedule that aligns effectively with the characteristics of the data. The spectral perspective also provides insights into the underlying dynamics and sheds light on the relationship between spectral properties and noise schedule structure. Our results lead to scheduling curves that are dependent on the spectral content of the data, offering a theoretical justification for some of the heuristics taken by practitioners.


Quadratic Coreset Selection: Certifying and Reconciling Sequence and Token Mining for Efficient Instruction Tuning

Neural Information Processing Systems

Instruction-Tuning (IT) was recently found the impressive data efficiency in posttraining large language models (LLMs). While the pursuit of efficiency predominantly focuses on sequence-level curation, often overlooking the nuanced impact of critical tokens and the inherent risks of token noise and biases. Drawing inspiration from bi-level coreset selection, our work provides the principled view of the motivation behind selecting instructions' responses. It leads to our approach Quadratic Coreset Selection (QCS) that reconciles sequence-level and token-level influence contributions, deriving more expressive LLMs with established theoretical result. Despite the original QCS framework challenged by prohibitive computation from inverted LLM-scale Hessian matrices, we overcome this barrier by proposing a novel QCS probabilistic variant, which relaxes the original formulation through re-parameterized densities. This innovative solver is efficiently learned using hierarchical policy gradients without requiring back-propagation, achieving provable convergence and certified asymptotic equivalence to the original objective. Our experiments demonstrate QCS's superior sequence-level data efficiency and reveal how strategically leveraging token-level influence elevates the performance ceiling of data-efficient IT. Furthermore, QCS's adaptability is showcased through its successes in regular IT and challenging targeted IT scenarios, particularly in the cases of free-form complex instruction-following and CoT reasoning. They underscore QCS's potential for a wide array of versatile post-training applications.


DepthVanish: Optimizing Adversarial Interval Structures for Stereo-Depth-Invisible Patches

Neural Information Processing Systems

Stereo depth estimation is a critical task in autonomous driving and robotics, where inaccuracies (such as misidentifying nearby objects as distant) can lead to dangerous situations. Adversarial attacks against stereo depth estimation can help reveal vulnerabilities before deployment. Previous works have shown that repeating optimized textures can effectively mislead stereo depth estimation in digital settings. However, our research reveals that these naively repeated textures perform poorly in physical implementations, i.e., when deployed as patches, limiting their practical utility for stress-testing stereo depth estimation systems. In this work, for the first time, we discover that introducing regular intervals among the repeated textures, creating a grid structure, significantly enhances the patch's attack performance. Through extensive experimentation, we analyze how variations of this novel structure influence the adversarial effectiveness. Based on these insights, we develop a novel stereo depth attack that jointly optimizes both the interval structure and texture elements. Our generated adversarial patches can be inserted into any scenes and successfully attack advanced stereo depth estimation methods of different paradigms, i.e., RAFT-Stereo and STTR. Most critically, our patch can also attack commercial RGB-D cameras (Intel RealSense) in real-world conditions, demonstrating their practical relevance for security assessment of stereo systems.


Let LRMs Break Free from Overthinking via Self-Braking Tuning

Neural Information Processing Systems

Large reasoning models (LRMs), such as OpenAI o1 and DeepSeek-R1, have significantly enhanced their reasoning capabilities by generating longer chains of thought, demonstrating outstanding performance across a variety of tasks. However, this performance gain comes at the cost of a substantial increase in redundant reasoning during the generation process, leading to high computational overhead and exacerbating the issue of overthinking. Although numerous existing approaches aim to address the problem of overthinking, they often rely on external interventions. In this paper, we propose a novel framework, Self-Braking Tuning (SBT), which tackles overthinking from the perspective of allowing the model to regulate its own reasoning process, thus eliminating the reliance on external control mechanisms. We construct a set of overthinking identification metrics based on standard answers and design a systematic method to detect redundant reasoning. This method accurately identifies unnecessary steps within the reasoning trajectory and generates training signals for learning self-regulation behaviors. Building on this foundation, we develop a complete strategy for constructing data with adaptive reasoning lengths and introduce an innovative braking prompt mechanism that enables the model to naturally learn when to terminate reasoning at an appropriate point. Experiments across mathematical benchmarks (AIME, AMC, MATH500, GSM8K) demonstrate that our method reduces token consumption by up to 60% while maintaining comparable accuracy to unconstrained models.


Global capitalism bets it all on AI future, alarming voters

The Japan Times

Days after filing confidentially to go public, Anthropic, the $965 billion artificial intelligence juggernaut that's one of the fastest-growing startups of all time, dropped another bombshell. In a blog post, Anthropic suggested the world might benefit from a slowdown in development of the very technologies that have been minting cash for the company. Provided global peers agreed, and enforcement mechanisms could be set up, that would help societies deal with the "immense implications" of AI, it said. Critics have long accused Anthropic of "doom marketing" -- hyping its own products as so good that they're bad. But the post's co-author, who's also the company's co-founder, says the motive is very different. "We say this stuff because we think the world needs to know the truth about what's happening," Jack Clark, who now heads Anthropic's public benefit work, said in an interview.


GTPBD: A Fine-Grained Global Terraced Parcel and Boundary Dataset

Neural Information Processing Systems

Agricultural parcels serve as basic units for conducting agricultural practices and applications, which is vital for land ownership registration, food security assessment, soil erosion monitoring, etc. However, existing agriculture parcel extraction studies only focus on mid-resolution mapping or regular plain farmlands while lacking representation of complex terraced terrains due to the demands of precision agriculture. In this paper, we introduce a more fine-grained terraced parcel dataset named GTPBD (Global Terraced Parcel and Boundary Dataset), which is the first fine-grained dataset covering major worldwide terraced regions with more than 200,000 complex terraced parcels with manually annotation. GTPBD comprises 47,537 high-resolution images with three-level labels, including pixel-level boundary labels, mask labels, and parcel labels. It covers seven major geographic zones in China and transcontinental climatic regions around the world. Compared to the existing datasets, the GTPBD dataset brings considerable challenges due to the: (1) terrain diversity; (2) complex and irregular parcel objects; and (3) multiple domain styles. Our proposed GTPBD dataset is suitable for four different tasks, including semantic segmentation, edge detection, terraced parcel extraction and unsupervised domain adaptation (UDA) tasks.


Pat McAfee wages war on Omaha's famous Jell-o shot bar after crew gets cold reception at College World Series

FOX News

NASCAR legend Tony Stewart calls mourning fans'a--holes' in tone-deaf rant about Kyle Busch Brewers' Jacob Misiorowski breaks brains and radar guns with hardest pitch ever by a starting pitcher US fans were out in full force ahead of the USMNT's first match of the 2026 FIFA World Cup MLB announces drive-in theater screenings of'The Sandlot' with live games and fireworks for July 4th California Democratic Party under fire for'you're not allowed to watch' World Cup post Victor Wembanyama isn't good or mature enough to be the face of the NBA -- at least not yet Rep. Byron Donalds shares his faith redemption story amid Florida gubernatorial run Iran's foreign minister says peace with US'has never been closer' GOP lawmaker says it's'really important' that US continues cartel crackdown Spencer Pratt's use of AI to boost campaign sparks debate FBI arrests first suspect on'most wanted fraudsters' list Accused Charlie Kirk killer's attorneys seek to BLOCK death penalty Kayleigh McEnany: Capitalism isn't the big evil Bernie Sanders would have you believe OutKick Sports Pat McAfee wages war on Omaha's famous Jell-o shot bar after crew gets cold reception at College World Series McAfee says the general manager was unhappy he didn't call ahead and mocked his ability to pay for shots Dan Dakich asks how ESPN's relevance has changed since adding Pat McAfee. We've got drama at the College World Series, and it has nothing to do with baseball. Pat McAfee has waged war with Rocco's -- the famous Omaha-based bar known for its Jell-O shot challenge during the 12-day tournament. And by war, I mean McAfee stuffed the GM in a locker during a heated segment on his ESPN and YouTube show Friday afternoon. It was nowhere near what I thought it was going to be like, McAfee said of the crew's experience at the bar earlier this week.