Large Language Model
Adaptive Preference Optimization with Uncertainty-aware Utility Anchor
Wang, Xiaobo, Jia, Zixia, Li, Jiaqi, Liu, Qi, Zheng, Zilong
Offline preference optimization methods are efficient for large language models (LLMs) alignment. Direct Preference optimization (DPO)-like learning, one of the most popular approaches, stands out for its efficiency in reward modeling. However, these methods typically follow the convention to use Bradley-Terry (BT) reward modeling that faces several critical assumptions, including the requirement for pairwise training data, model distribution shifting, human rationality assumption, etc. To address these limitations, we propose a general framework for offline preference optimization methods, Adaptive Preference Optimization with Utility Anchor (UAPO), which introduces an anchoring function to estimate the uncertainties brought from preference data annotation. Our method enables training even in scenarios where the data is unpaired, significantly enhancing data utilization efficiency. Moreover, the anchor design makes UAPO more robust in the training process. Experimental results demonstrate that UAPO achieves competitive outcomes without the strict dependency on data pairing, paving the way for more flexible and effective preference optimization methods.
Mixture-of-Clustered-Experts: Advancing Expert Specialization and Generalization in Instruction Tuning
Eo, Sugyeong, Lee, Jungjun, Park, Chanjun, Lim, Heuiseok
A sparse Mixture-of-Experts (MoE) architecture has emerged as a highly scalable solution by conditionally activating sub-modules without a proportional increase in computational costs. However, improving expert specialization to enhance performance and generalization remains a challenge for MoE, especially in instruction tuning scenarios characterized by significant input heterogeneity. In this work, we propose the Mixture-of-Clustered-Experts (MoCE) to address this limitation through a dual-stage routing mechanism. The first stage in the mechanism performs expert group routing based on sequence-level features, while the second stage activates the top-$k$ experts within the group at the token level. This approach enables the effective partitioning of heterogeneous inputs based on their knowledge requirements, encouraging expert group specialization while maintaining the advantages of token-level routing. We evaluate MoCE across a comprehensive set of benchmarks, demonstrating its consistent superiority over strong baselines and its enhanced generalization capabilities. Detailed analysis further highlights the robustness and effectiveness of MoCE.
The Anti-Ouroboros Effect: Emergent Resilience in Large Language Models from Recursive Selective Feedback
B. Experiment 2: The Anti-Ouroboros Effect in an LLM In the LLM experiment, the results falsified the hypothesis. The Quality Filter arm demonstrated robust and statistically significant improvement across all three evaluation metrics, as shown in Table II. In contrast, both the unfiltered Control arm and the Random Filter arm exhibited performance degradation, proving that the improvement is due to intelligent selection, not merely training on less data. The performance trajectories for all three arms are visualized in Figure 1 C. Human Evaluation To provide an independent verification of our automated metrics, we conducted a small, blinded human study. Two evaluators rated 30 anonymized and shuffled summaries from the final generation of the Control and Quality Filter arms on a 1-5 scale. The Quality Filter arm significantly outperformed the Control arm on coherence (4.2 vs 3.5) and factuality (4.5 vs 3.8), confirming the quantitative results. D. Analysis of the Mechanism The emergence of the Anti-Ouroboros Effect in the LLM suggests a dynamic unique to high-dimensional systems. We propose two non-exclusive hypotheses: Error Propagation Shutdown, where the filter acts as a ratchet, preventing the reinforcement of errors, and Latent Space Guidance, where selection guides the fine-tuning process toward more robust regions of the model's parameter space.
The LLM as a Network Operator: A Vision for Generative AI in the 6G Radio Access Network
Giwa, Oluwaseyi, Adewole, Michael, Awodumila, Tobi, Aderinto, Pelumi
The management of future AI-native Next-Generation (NextG) Radio Access Networks (RANs), including 6G and beyond, presents a challenge of immense complexity that exceeds the capabilities of traditional automation. In response, we introduce the concept of the LLM-RAN Operator. In this paradigm, a Large Language Model (LLM) is embedded into the RAN control loop to translate high-level human intents into optimal network actions. Unlike prior empirical studies, we present a formal framework for an LLM-RAN operator that builds on earlier work by making guarantees checkable through an adapter aligned with the Open RAN (O-RAN) standard, separating strategic LLM-driven guidance in the Non-Real-Time (RT) RAN intelligent controller (RIC) from reactive execution in the Near-RT RIC, including a proposition on policy expressiveness and a theorem on convergence to stable fixed points. By framing the problem with mathematical rigor, our work provides the analytical tools to reason about the feasibility and stability of AI-native RAN control. It identifies critical research challenges in safety, real-time performance, and physical-world grounding. This paper aims to bridge the gap between AI theory and wireless systems engineering in the NextG era, aligning with the AI4NextG vision to develop knowledgeable, intent-driven wireless networks that integrate generative AI into the heart of the RAN.
Learning Decomposed Contextual Token Representations from Pretrained and Collaborative Signals for Generative Recommendation
Liu, Yifan, Liu, Yaokun, Li, Zelin, Yue, Zhenrui, Lee, Gyuseok, Yao, Ruichen, Zhang, Yang, Wang, Dong
Recent advances in generative recommenders adopt a two-stage paradigm: items are first tokenized into semantic IDs using a pretrained tokenizer, and then large language models (LLMs) are trained to generate the next item via sequence-to-sequence modeling. However, these two stages are optimized for different objectives: semantic reconstruction during tokenizer pretraining versus user interaction modeling during recommender training. This objective misalignment leads to two key limitations: (i) suboptimal static tokeniza-tion, where fixed token assignments fail to reflect diverse usage contexts; and (ii) discarded pretrained semantics, where pretrained knowledge--typically from language model em-beddings--is overwritten during recommender training on user interactions. To address these limitations, we propose to learn DE composed CO ntextual Token R epresentations (DECOR), a unified framework that preserves pretrained semantics while enhancing the adaptability of token embed-dings. DECOR introduces contextualized token composition to refine token embeddings based on user interaction context, and decomposed embedding fusion that integrates pretrained codebook embeddings with newly learned collaborative em-beddings. Experiments on three real-world datasets demonstrate that DECOR consistently outperforms state-of-the-art baselines in recommendation performance. Our code will be made available upon publication.
DSRAG: A Domain-Specific Retrieval Framework Based on Document-derived Multimodal Knowledge Graph
Yang, Mengzheng, Ren, Yanfei, Opoku, David Osei, Li, Ruochang, Ren, Peng, Xing, Chunxiao
Retrieval-augmented generation (RAG) effectively tackles these challenges by integrating external knowledge to enhance accuracy and relevance. However, traditional RAG still faces limitations in domain knowledge accuracy and context modeling.To enhance domain-specific question answering performance, this work focuses on a graph-based RAG framework, emphasizing the critical role of knowledge graph quality during the generation process. We propose DSRAG (Domain-Specific RAG), a multimodal knowledge graph-driven retrieval-augmented generation framework designed for domain-specific applications. Our approach leverages domain-specific documents as the primary knowledge source, integrating heterogeneous information such as text, images, and tables to construct a multimodal knowledge graph covering both conceptual and instance layers. Building on this foundation, we introduce semantic pruning and structured subgraph retrieval mechanisms, combining knowledge graph context and vector retrieval results to guide the language model towards producing more reliable responses.
The 1st International Workshop on Disentangled Representation Learning for Controllable Generation (DRL4Real): Methods and Results
Chen, Qiuyu, Jin, Xin, Song, Yue, Liu, Xihui, Yang, Shuai, Yang, Tao, Li, Ziqiang, Huang, Jianguo, Wei, Yuntao, Xie, Ba'ao, Sebe, Nicu, Wenjun, null, Zeng, null, Yun, Jooyeol, Abati, Davide, Omran, Mohamed, Choo, Jaegul, Habibian, Amir, Wiggers, Auke, Kobayashi, Masato, Ding, Ning, Tamaki, Toru, Gheisari, Marzieh, Genovesio, Auguste, Chen, Yuheng, Liu, Dingkun, Yang, Xinyao, Xu, Xinping, Chen, Baicheng, Wu, Dongrui, Geng, Junhao, Lv, Lexiang, Lin, Jianxin, Liang, Hanzhe, Zhou, Jie, Chen, Xuanxin, Wang, Jinbao, Gao, Can, Wang, Zhangyi, Li, Zongze, Wen, Bihan, Gao, Yixin, Pan, Xiaohan, Li, Xin, Chen, Zhibo, Peng, Baorui, Chen, Zhongming, Jin, Haoran
This paper reviews the 1st International Workshop on Disentangled Representation Learning for Controllable Generation (DRL4Real), held in conjunction with ICCV 2025. The workshop aimed to bridge the gap between the theoretical promise of Disentangled Representation Learning (DRL) and its application in realistic scenarios, moving beyond synthetic benchmarks. DRL4Real focused on evaluating DRL methods in practical applications such as controllable generation, exploring advancements in model robustness, interpretability, and generalization. The workshop accepted 9 papers covering a broad range of topics, including the integration of novel inductive biases (e.g., language), the application of diffusion models to DRL, 3D-aware disentanglement, and the expansion of DRL into specialized domains like autonomous driving and EEG analysis. This summary details the workshop's objectives, the themes of the accepted papers, and provides an overview of the methodologies proposed by the authors.
Towards Understanding Visual Grounding in Visual Language Models
Pantazopoulos, Georgios, รzyiฤit, Eda B.
Visual grounding refers to the ability of a model to identify a region within some visual input that matches a textual description. Consequently, a model equipped with visual grounding capabilities can target a wide range of applications in various domains, including referring expression comprehension, answering questions pertinent to fine-grained details in images or videos, caption visual context by explicitly referring to entities, as well as low and high-level control in simulated and real environments. In this survey paper, we review representative works across the key areas of research on modern general-purpose vision language models (VLMs). We first outline the importance of grounding in VLMs, then delineate the core components of the contemporary paradigm for developing grounded models, and examine their practical applications, including benchmarks and evaluation metrics for grounded multimodal generation. We also discuss the multifaceted interrelations among visual grounding, multimodal chain-of-thought, and reasoning in VLMs. Finally, we analyse the challenges inherent to visual grounding and suggest promising directions for future research.
Towards Reliable and Interpretable Document Question Answering via VLMs
Chen, Alessio, Giovannini, Simone, Gemelli, Andrea, Coppini, Fabio, Marinai, Simone
Vision-Language Models (VLMs) have shown strong capabilities in document understanding, particularly in identifying and extracting textual information from complex documents. Despite this, accurately localizing answers within documents remains a major challenge, limiting both interpretability and real-world applicability. T o address this, we introduce DocExplainerV0, a plug-and-play bounding-box prediction module that decouples answer generation from spatial localization. This design makes it applicable to existing VLMs, including proprietary systems where fine-tuning is not feasible. Through systematic evaluation, we provide quantitative insights into the gap between textual accuracy and spatial grounding, showing that correct answers often lack reliable localization. Our standardized framework highlights these shortcomings and establishes a benchmark for future research toward more interpretable and robust document information extraction VLMs.
GeoGPT-RAG Technical Report
Huang, Fei, Wu, Fan, Zhang, Zeqing, Wang, Qihao, Zhang, Long, Boquet, Grant Michael, Chen, Hongyang
GeoGPT is an open large language model system built to advance research in the geosciences. To enhance its domain-specific capabilities, we integrated Retrieval Augmented Generation(RAG), which augments model outputs with relevant information retrieved from an external knowledge source. GeoGPT uses RAG to draw from the GeoGPT Library, a specialized corpus curated for geoscientific content, enabling it to generate accurate, context-specific answers. Users can also create personalized knowledge bases by uploading their own publication lists, allowing GeoGPT to retrieve and respond using user-provided materials. To further improve retrieval quality and domain alignment, we fine-tuned both the embedding model and a ranking model that scores retrieved passages by relevance to the query. These enhancements optimize RAG for geoscience applications and significantly improve the system's ability to deliver precise and trustworthy outputs. GeoGPT reflects a strong commitment to open science through its emphasis on collaboration, transparency, and community driven development. As part of this commitment, we have open-sourced two core RAG components-GeoEmbedding and GeoReranker-to support geoscientists, researchers, and professionals worldwide with powerful, accessible AI tools.