Goto

Collaborating Authors

 Media


Visual Discovering Object Dependencies via Counterfactual

Neural Information Processing Systems

This paper proposes a novel scene understanding task called Visual Jenga. Drawing inspiration from the game Jenga, the proposed task involves progressively removing objects from a single image until only the background remains. Just as Jenga players must understand structural dependencies to maintain tower stability, our task reveals the intrinsic relationships between scene elements by systematically exploring which objects can be removed while preserving scene coherence in both physical and geometric sense. As a starting point for tackling the Visual Jenga task, we propose a simple, data-driven, training-free approach that is surprisingly effective on a range of real-world images. The principle behind our approach is to utilize the asymmetry in the pairwise relationships between objects within a scene and employ a large inpainting model to generate a set of counterfactuals to quantify the asymmetry.


DisMo: Disentangled Motion Representations for Open-World Motion Transfer

Neural Information Processing Systems

Recent advances in text-to-video (T2V) and image-to-video (I2V) models, have enabled the creation of visually compelling and dynamic videos from simple textual descriptions or initial frames. However, these models often fail to provide an explicit representation of motion separate from content, limiting their applicability for content creators. To address this gap, we propose DisMo, a novel paradigm for learning abstract motion representations directly from raw video data via an image-space reconstruction objective. Our representation is generic and independent of static information such as appearance, object identity, or pose. This enables open-world motion transfer, allowing motion to be transferred across semantically unrelated entities without requiring object correspondences, even between vastly different categories. Unlike prior methods, which trade off motion fidelity and prompt adherence, are overfitting to source structure or drifting from the described action, our approach disentangles motion semantics from appearance, enabling accurate transfer and faithful conditioning. Furthermore, our motion representation can be combined with any existing video generator via lightweight adapters, allowing us to effortlessly benefit from future advancements in video models. We demonstrate the effectiveness of our method through a diverse set of motion transfer tasks. Finally, we show that the learned representations are well-suited for downstream motion understanding tasks, consistently outperforming state-of-the-art video representation models such as V-JEPA in zero-shot action classification on benchmarks including Something-Something v2 and Jester.



KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models

Neural Information Processing Systems

Recent advances in multi-modal generative models have enabled significant progress in instruction-based image editing. However, while these models produce visually plausible outputs, their capacity for knowledge-based reasoning editing tasks remains under-explored. In this paper, We introduce KRIS-Bench (Knowledge-based Reasoning in Image-editing Systems Benchmark), a diagnostic benchmark designed to assess models through a cognitively informed lens.


Non-Markovian Discrete Diffusion with Causal Language Models

Neural Information Processing Systems

Discrete diffusion models offer a flexible, controllable approach to structured sequence generation, yet they still lag behind causal language models in expressive power. A key limitation lies in their reliance on the Markovian assumption, which restricts each step to condition only on the current state, leading to potential uncorrectable error accumulation. In this paper, we introduce CaDDi (Causal Discrete Diffusion Model), a discrete diffusion model that conditions on the entire generative trajectory, thereby lifting the Markov constraint and allowing the model to revisit and improve past states. By unifying sequential (causal) and temporal (diffusion) reasoning in a single non-Markovian transformer, CaDDi also treats standard causal language models as a special case and permits the direct reuse of pretrained LLM weights with no architectural changes. Empirically, CaDDi outperforms state-of-the-art discrete diffusion baselines on natural-language benchmarks, substantially narrowing the remaining gap to large autoregressive transformers.


Removing Concepts from Text-to-Image Models with Only Negative Samples

Neural Information Processing Systems

This work introduces Clipout, a method for removing a target concept in pretrained text-to-image models. By randomly clipping units from the learned data embedding and using a contrastive objective, models are encouraged to differentiate these clipped embedding vectors.


Distance infer View change infer More tasks

Neural Information Processing Systems

Instead of injecting 3D representations, we unlock VLMs using spatially relevant 2D images. To this end, we introduce a novel 2D spatial data generation and enables annotation the creation pipeline of a b di uilt verse upon set scene of spatial data tasks, with 3D ranging ground-truth.


e36da7acd188c6655792270b38830124-Paper-Conference.pdf

Neural Information Processing Systems

Diffusion Transformers (DiT) have become a leading architecture in image generation. However, the quadratic complexity of attention mechanisms, which are when responsible generating for modeling high-resolution token-wis images.


Concept Incongruence: An Exploration of Time and Death in Role Playing

Neural Information Processing Systems

Consider this prompt "Draw a unicorn with two horns". Should large language models (LLMs) recognize that a unicorn has only one horn by definition and ask users for clarifications, or proceed to generate something anyway? We introduce concept incongruence to capture such phenomena where concept boundaries clash with each other, either in user prompts or in model representations, often leading to under-specified or mis-specified behaviors. In this work, we take the first step towards defining and analyzing model behavior under concept incongruence. Focusing on temporal boundaries in the ROLE-PLAY setting, we propose three behavioral metrics--abstention rate, conditional accuracy, and answer rate--to quantify model behavior under incongruence due to the role's death. We show that models fail to abstain after death and suffer from an accuracy drop compared to the NON-ROLE-PLAY setting. Through probing experiments, we identify two main causes: (i) unreliable encoding of the "death" state across different years, leading to unsatisfactory abstention behavior, and (ii) role playing causes shifts in the model's temporal representations, resulting in accuracy drops. We leverage these insights to improve consistency in the model's abstention and answer behaviors. Our findings suggest that concept incongruence leads to unexpected model behaviors and point to future directions on improving model behavior under concept incongruence.1


e1ebda145808ca45774993fb67314894-Supplemental-Datasets_and_Benchmarks_Track.pdf

Neural Information Processing Systems

ARelated Work1 Data Attribution Evaluation: Given recent developments in data attribution methods for LLMs,2 past works in evaluating these methods fall two major categories: leave-out-out and task-based3 evaluation. Leave-one-out evaluation measures the correlation between the data attribution method4 scores and model-retraining, which can also be approximated using linear datamodeling score [26].5 In task-based evaluation, the data attribution method is evaluated based on its application towards6 downstream task, such as noisy label detection, counterfactual evaluation [3, 13].7 Training Data Selection: Selecting high-quality training data selection is important for efficient8 learning in LLMs. Common approaches to data selection relies on heuristic filtering, such as de-9 duplication and lexicon-filtering, [34], or semantic rating [48, 52]. Recent works have applied data10 attribution methods towards data selection in LLMs in both pre-training [56, 59, 15] and post-training11 [45, 53, 31]. These data attribution methods are dynamic and model-aware - increasing the frequency12 of performing selection is one way to take greater account for group influence, where online selection13 at each training step is most fine-grained [49].14 Toxicity/Bias Detection: Detecting and mitigating toxic/biased LLMs outputs is a crucial for safe15 deployment in real-word settings. Existing methods for detecting toxicity/bias in LLMs commonly16 include online API tools 1 [37] or LLM-classifiers [58, 21, 16, 27]. Factual Attribution: Identifying training examples which causes LLMs to generate specific factual20 statements is an important application of data attribution as AI tools are becoming increasingly21 common. Apart from baseline retrieval methods that leverage lexical/semantic similarity like BM2522 [48], Rep Sim [44] and Gecko [33], recent works have explored the use of data attribution in tracing23 factual knowledge in both pre-training[6] and post-training [42, 2].24 We provide below descriptions to the data attribution methods and non-attribution baselines evaluated26 in this work. Note that in our work, we consider non-attribution baselines as methods that do not27 estimate the impact of training samples on models, as detailed in [19].28 Rep-Sim [44]: (Non-attribution baseline) Rep-Sim computes the cosine similarity between last29 token last layer hidden states of training and reference examples. It is more efficient compared with30 gradient-based data attribution methods. BM25 [48]: (Non-attribution baseline) BM25 is a classic information retrieval algorithm that ranks33 training samples by lexical overlap with the query. It is significantly more efficient compared with34 gradient-based data attribution methods.35