Deep Learning
The Curious Case of Visual Grounding: Different Effects for Speech- and Text-based Language Encoders
Sauter, Adrian, Zuidema, Willem, Kloots, Marianne de Heer
How does visual information included in training affect language processing in audio- and text-based deep learning models? We explore how such visual grounding affects model-internal representations of words, and find substantially different effects in speech- vs. text-based language encoders. Firstly, global representational comparisons reveal that visual grounding increases alignment between representations of spoken and written language, but this effect seems mainly driven by enhanced encoding of word identity rather than meaning. We then apply targeted clustering analyses to probe for phonetic vs. semantic discriminability in model representations. Speech-based representations remain phonetically dominated with visual grounding, but in contrast to text-based representations, visual grounding does not improve semantic discriminability. Our findings could usefully inform the development of more efficient methods to enrich speech-based models with visually-informed semantics.
CIDER: A Causal Cure for Brand-Obsessed Text-to-Image Models
Shen, Fangjian, Liang, Zifeng, Wang, Chao, Wen, Wushao
Text-to-image (T2I) models exhibit a significant yet under-explored "brand bias", a tendency to generate contents featuring dominant commercial brands from generic prompts, posing ethical and legal risks. We propose CIDER, a novel, model-agnostic framework to mitigate bias at inference-time through prompt refinement to avoid costly retraining. CIDER uses a lightweight detector to identify branded content and a Vision-Language Model (VLM) to generate stylistically divergent alternatives. We introduce the Brand Neutrality Score (BNS) to quantify this issue and perform extensive experiments on leading T2I models. Results show CIDER significantly reduces both explicit and implicit biases while maintaining image quality and aesthetic appeal. Our work offers a practical solution for more original and equitable content, contributing to the development of trustworthy generative AI.
ChronoForge-RL: Chronological Forging through Reinforcement Learning for Enhanced Video Understanding
In this paper, we propose a novel video understanding framework, called ChronoForge-RL, which combines Temporal Apex Distillation (T AD) and KeyFrame-aware Group Relative Policy Optimization (KF-GRPO) to tackle these issues. Concretely, we introduce a differentiable keyframe selection mechanism that systematically identifies semantic inflection points through a three-stage process to enhance computational efficiency while preserving temporal information. Then, two particular modules are proposed to enable effective temporal reasoning: Firstly, T AD leverages variation scoring, inflection detection, and prioritized distillation to select the most informative frames. Secondly, we introduce KF-GRPO which implements a contrastive learning paradigm with a saliency-enhanced reward mechanism that explicitly incentivizes models to leverage both frame content and temporal relationships. Finally, our proposed ChronoForge-RL achieves 69.1% on VideoMME and 52.7% on L VBench compared to baseline methods, clearly surpassing previous approaches while enabling our 7B parameter model to achieve performance comparable to 72B parameter alternatives, a 10 improvement in performance-to-parameter ratio.
Monte Carlo Tree Diffusion with Multiple Experts for Protein Design
Liu, Xuefeng, Cao, Mingxuan, Jiang, Songhao, Luo, Xiao, Duan, Xiaotian, Wang, Mengdi, Sosnick, Tobin R., Xu, Jinbo, Stevens, Rick
The goal of protein design is to generate amino acid sequences that fold into functional structures with desired properties. Prior methods combining autoregressive language models with Monte Carlo Tree Search (MCTS) struggle with long-range dependencies and suffer from an impractically large search space. We propose MCTD-ME, Monte Carlo Tree Diffusion with Multiple Experts, which integrates masked diffusion models with tree search to enable multi-token planning and efficient exploration. Unlike autoregressive planners, MCTD-ME uses biophysical-fidelity-enhanced diffusion denoising as the rollout engine, jointly revising multiple positions and scaling to large sequence spaces. It further leverages experts of varying capacities to enrich exploration, guided by a pLDDT-based masking schedule that targets low-confidence regions while preserving reliable residues. We propose a novel multi-expert selection rule (PH-UCT-ME) extends predictive-entropy UCT to expert ensembles. On the inverse folding task (CAMEO and PDB benchmarks), MCTD-ME outperforms single-expert and unguided baselines in both sequence recovery (AAR) and structural similarity (scTM), with gains increasing for longer proteins and benefiting from multi-expert guidance. More generally, the framework is model-agnostic and applicable beyond inverse folding, including de novo protein engineering and multi-objective molecular generation.
Building Data-Driven Occupation Taxonomies: A Bottom-Up Multi-Stage Approach via Semantic Clustering and Multi-Agent Collaboration
Li, Nan, Kang, Bo, De Bie, Tijl
Creating robust occupation taxonomies, vital for applications ranging from job recommendation to labor market intelligence, is challenging. Manual curation is slow, while existing automated methods are either not adaptive to dynamic regional markets (top-down) or struggle to build coherent hierarchies from noisy data (bottom-up). We introduce CLIMB (CLusterIng-based Multi-agent taxonomy Builder), a framework that fully automates the creation of high-quality, data-driven taxonomies from raw job postings. CLIMB uses global semantic clustering to distill core occupations, then employs a reflection-based multi-agent system to iteratively build a coherent hierarchy. On three diverse, real-world datasets, we show that CLIMB produces taxonomies that are more coherent and scalable than existing methods and successfully capture unique regional characteristics. We release our code and datasets at https://anonymous.4open.science/r/CLIMB.
Ideal Registration? Segmentation is All You Need
Chen, Xiang, Zhang, Fengting, Liu, Qinghao, Liu, Min, Wu, Kun, Wang, Yaonan, Zhang, Hang
Deep learning has revolutionized image registration by its ability to handle diverse tasks while achieving significant speed advantages over conventional approaches. Current approaches, however, often employ globally uniform smoothness constraints that fail to accommodate the complex, regionally varying deformations characteristic of anatomical motion. To address this limitation, we propose SegReg, a Segmentation-driven Registration framework that implements anatomically adaptive regularization by exploiting region-specific deformation patterns. Our SegReg first decomposes input moving and fixed images into anatomically coherent subregions through segmentation. These localized domains are then processed by the same registration backbone to compute optimized partial deformation fields, which are subsequently integrated into a global deformation field. SegReg achieves near-perfect structural alignment (98.23% Dice on critical anatomies) using ground-truth segmentation, and outperforms existing methods by 2-12% across three clinical registration scenarios (cardiac, abdominal, and lung images) even with automatic segmentation. Our SegReg demonstrates a near-linear dependence of registration accuracy on segmentation quality, transforming the registration challenge into a segmentation problem. The source code will be released upon manuscript acceptance.
UniGist: Towards General and Hardware-aligned Sequence-level Long Context Compression
Deng, Chenlong, Zhang, Zhisong, Mao, Kelong, Li, Shuaiyi, Fang, Tianqing, Zhang, Hongming, Mi, Haitao, Yu, Dong, Dou, Zhicheng
Large language models are increasingly capable of handling long-context inputs, but the memory overhead of key-value (KV) cache remains a major bottleneck for general-purpose deployment. While various compression strategies have been explored, sequence-level compression, which drops the full KV caches for certain tokens, is particularly challenging as it can lead to the loss of important contextual information. To address this, we introduce UniGist, a sequence-level long-context compression framework that efficiently preserves context information by replacing raw tokens with special compression tokens (gists) in a fine-grained manner. We adopt a chunk-free training strategy and design an efficient kernel with a gist shift trick, enabling optimized GPU training. Our scheme also supports flexible inference by allowing the actual removal of compressed tokens, resulting in real-time memory savings. Experiments across multiple long-context tasks demonstrate that UniGist significantly improves compression quality, with especially strong performance in detail-recalling tasks and long-range dependency modeling.
FloorSAM: SAM-Guided Floorplan Reconstruction with Semantic-Geometric Fusion
Ye, Han, Wang, Haofu, Zhang, Yunchi, Xiao, Jiangjian, Jin, Yuqiang, Liu, Jinyuan, Zhang, Wen-An, Sychou, Uladzislau, Tuzikov, Alexander, Sobolevskii, Vladislav, Zakharov, Valerii, Sokolov, Boris, Fu, Minglei
Abstract--Reconstructing building floor plans from point cloud data is a critical technology for indoor navigation, building information modeling (BIM), and highly accurate precise indoor measurement applications. Traditional methods, such as geometric algorithms and Mask R-CNN-based deep learning for mask segmentation, often suffer from sensitivity to noise, limited generalization, and loss of geometric details, severely impacting measurement accuracy. This study proposes an innovative framework, FloorSAM, that integrates room-height point cloud density maps with the guided segmentation capabilities of the Segment Anything Model (SAM) to enhance the precision of floor plan reconstruction from LiDAR point cloud data. By applying grid-based filtering to retain elevation point clouds near the ceiling of each region, combined with adaptive resolution projection and image enhancement techniques, a top-down density map is generated, improving the robustness and accuracy of spatial feature measurement. This framework leverages SAM's zero-shot learning to achieve high-fidelity room segmentation, remarkably enhancing reconstruction and measurement accuracy across diverse building layouts. Subsequently, leveraging SAM's zero-shot guided segmentation capabilities, high-quality room masks are generated based on adaptive prompt points, followed by a multistage filtering process to extract precise semantic masks for individual rooms. Through joint analysis of mask and point cloud modalities, contour extraction and regularization are performed, integrating semantic segmentation with geometric information to produce accurate room floor plans and recover topological relationships between rooms.
Can LLMs Judge Debates? Evaluating Non-Linear Reasoning via Argumentation Theory Semantics
Sanayei, Reza, Vesic, Srdjan, Blanco, Eduardo, Surdeanu, Mihai
Large Language Models (LLMs) excel at linear reasoning tasks but remain underexplored on non-linear structures such as those found in natural debates, which are best expressed as argument graphs. We evaluate whether LLMs can approximate structured reasoning from Computational Argumentation Theory (CAT). Specifically, we use Quantitative Argumentation Debate (QuAD) semantics, which assigns acceptability scores to arguments based on their attack and support relations. Given only dialogue-formatted debates from two NoDE datasets, models are prompted to rank arguments without access to the underlying graph. We test several LLMs under advanced instruction strategies, including Chain-of-Thought and In-Context Learning. While models show moderate alignment with QuAD rankings, performance degrades with longer inputs or disrupted discourse flow. Advanced prompting helps mitigate these effects by reducing biases related to argument length and position. Our findings highlight both the promise and limitations of LLMs in modeling formal argumentation semantics and motivate future work on graph-aware reasoning.
A Nascent Taxonomy of Machine Learning in Intelligent Robotic Process Automation
Laakmann, Lukas, Ciftci, Seyyid A., Janiesch, Christian
Robotic process automation (RPA) is a lightweight approach to automating business processes using software robots that emulate user actions at the graphical user interface level. While RPA has gained popularity for its cost-effective and timely automation of rule-based, well-structured tasks, its symbolic nature has inherent limitations when approaching more complex tasks currently performed by human agents. Machine learning concepts enabling intelligent RPA provide an opportunity to broaden the range of automatable tasks. In this paper, we conduct a literature review to explore the connections between RPA and machine learning and organize the joint concept intelligent RPA into a taxonomy. Our taxonomy comprises the two meta-characteristics RPA-ML integration and RPA-ML interaction. Together, they comprise eight dimensions: architecture and ecosystem, capabilities, data basis, intelligence level, and technical depth of integration as well as deployment environment, lifecycle phase, and user-robot relation.