ProxyThinker: Test-Time Guidance through Small Visual Reasoners
Xiao, Zilin, Koo, Jaywon, Ouyang, Siru, Hernandez, Jefferson, Meng, Yu, Ordonez, Vicente
–arXiv.org Artificial Intelligence
Recent advancements in reinforcement learning with verifiable rewards have pushed the boundaries of the visual reasoning capabilities in large vision-language models (L VLMs). However, training L VLMs with reinforcement fine-tuning (RFT) is computationally expensive, posing a significant challenge to scaling model size. Recent advances in large language models have led to the development of systems capable of extended reasoning and deliberation, often referred to as "slow-thinking" models, such as OpenAI-o1 (Jaech et al., 2024), DeepSeek-R1 (Guo et al., 2025), and QwQ (Team, 2025). Unlike "fast-thinking" models such as GPT -4o (Hurst et al., 2024), "slow-thinking" models usually engage in multi-step self-reflection to produce an answer that resembles the thorough thinking process that humans make before producing a final answer for non-trivial problems. These models have achieved remarkable success in complex problem-solving benchmarks, particularly in mathematical and scientific reasoning domains (Shao et al., 2024b; Zeng et al., 2025; Y u et al., 2025). Recent research has also extended such reflective reasoning to multimodal tasks (Huang et al., 2025; Deng et al., 2025; Y ang et al., 2025; Zhou et al., 2025; Wang et al., 2025a), pushing large vision-language models (L VLMs) toward greater performance in scenarios that require structured and contextual understanding across modalities. Many of the most effective "slow-thinking" models rely on reinforcement learning with verifiable rewards (RL VR) (Face, 2025; Su et al., 2025; Wei et al., 2025a), a reinforcement fine-tuning (RFT) framework that encourages the model to generate intermediate reasoning steps that lead to a correct answer for automatically verifiable tasks.
arXiv.org Artificial Intelligence
Sep-30-2025
- Country:
- North America > United States (0.68)
- Europe > Austria
- Vienna (0.14)
- Genre:
- Research Report (1.00)
- Technology: