Mitigating Visual Hallucinations via Semantic Curriculum Preference Optimization in MLLMs

Li, Yuanshuai, Yan, Yuping, Tang, Junfeng, Li, Yunxuan, Zheng, Zeqi, Jin, Yaochu

arXiv.org Artificial Intelligence 

Multimodal Large Language Models (MLLMs) have significantly improved the performance of various tasks, but continue to suffer from visual hallucinations, a critical issue where generated responses contradict visual evidence. While Direct Preference Optimization (DPO) is widely used for alignment, its application to MLLMs often fails to capture fine-grained semantic differences and encourages shortcut learning. To address these challenges, we propose Semantic Curriculum Preference Optimization (SCPO), a novel framework for MLLM alignment. SCPO employs a progressive, easy-to-hard curriculum built upon our Semantic Curriculum Preference Pairs dataset, which provides fine-grained semantic contrasts sorted by difficulty. This curriculum is trained with a dynamic reference model and a novel symmetric, bidirectional objective to facilitate simultaneous learning from both textual and visual preferences. To our knowledge, SCPO is the first framework to unify semantics, symmetry, and curriculum for MLLMs alignment, effectively mitigating visual hallucinations. Extensive experiments on LLaV A models across various scales and versions validate that SCPO demonstrates superior performance compared to baseline models on multiple hallucination benchmarks, reducing the hallucination rate by up to 62.9%. Moreover, evaluations on generalized benchmarks show that SCPO improves factuality while preserving general capabilities, with its performance remaining stable across general vision-language benchmarks. Multimodal large language models (MLLMs) (Liu et al., 2024a; OpenAI, 2023; Liu et al., 2024b) have demonstrated remarkable capabilities in tasks requiring fine-grained perception, such as visual question answering and image captioning (Y u et al., 2025). Despite these advances, their reliability is compromised by the persistent challenge of visual hallucination. During inference, MLLMs often rely too heavily on the semantic priors embedded in the language model, while failing to faithfully ground these priors in fine-grained visual evidence from the input image (Li et al., 2022; Kalai et al., 2025; Xu et al., 2024b). This misalignment between language and vision produces outputs that contradict the actual visual content, thereby seriously limiting their applicability in high-stakes scenarios such as autonomous driving and medical image diagnosis (Y u et al., 2024a).

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found