Med-VRAgent: A Framework for Medical Visual Reasoning-Enhanced Agents
Guo, Guangfu, Lu, Xiaoqian, Feng, Yue
–arXiv.org Artificial Intelligence
Visual Language Models (VLMs) achieve promising results in medical reasoning but struggle with hallucinations, vague descriptions, inconsistent logic and poor localization. To address this, we propose a agent framework named Medical Visual Reasoning Agent (\textbf{Med-VRAgent}). The approach is based on Visual Guidance and Self-Reward paradigms and Monte Carlo Tree Search (MCTS). By combining the Visual Guidance with tree search, Med-VRAgent improves the medical visual reasoning capabilities of VLMs. We use the trajectories collected by Med-VRAgent as feedback to further improve the performance by fine-tuning the VLMs with the proximal policy optimization (PPO) objective. Experiments on multiple medical VQA benchmarks demonstrate that our method outperforms existing approaches.
arXiv.org Artificial Intelligence
Oct-22-2025
- Country:
- Europe > Spain
- Catalonia > Barcelona Province > Barcelona (0.04)
- North America > United States
- Michigan > Washtenaw County
- Ann Arbor (0.04)
- Pennsylvania > Philadelphia County
- Philadelphia (0.04)
- Michigan > Washtenaw County
- Europe > Spain
- Genre:
- Research Report > New Finding (0.46)
- Industry:
- Health & Medicine > Diagnostic Medicine > Imaging (0.48)
- Technology: