cs.AIOct 4, 2026

When to Rethink: Learning Multi-Perspective Self-Verification for Vision-Language Models

Authors: Ziquan Zhu, Hanruo Zhu, Si-Yuan Lu, Morris Yu-Chao Huang, Yicheng Lin, Wei Han, Tianlong Chen, Mingyuan Wu, +5 more

Abstract

Vision-language models (VLMs) have achieved strong performance in multimodal reasoning, yet they remain prone to generating plausible but incorrect answers. Self-verification offers a practical way to improve answer reliability without relying on external judges, but existing methods typically depend on a single verification criterion or fixed prompt, resulting in incomplete and unstable reliability estimates. We first systematically analyze how verifier capability and prompt design affect verification performance. Our findings show that stronger verifiers provide more reliable judgments, while verification performance is highly sensitive to prompt choice, with no single prompt consistently dominating across tasks. Guided by these findings, we propose \texttt{MOTIVE}, a \textbf{M}ulti-View Self-Verificati\textbf{O}n wi\textbf{T}h Rel\textbf{I}ability-Guided Selecti\textbf{VE} Rethinking framework for reliable multimodal reasoning. \texttt{MOTIVE} evaluates each candidate answer from complementary verification perspectives and learns a correctness-aligned reliability score through correctness-grounded multi-view verification learning. During inference, this score governs an accept-or-rethink decision, allowing reliable answers to be returned directly while uncertain ones trigger history-guided rethinking. Extensive experiments across diverse multimodal benchmarks and VLM backbones demonstrate that \texttt{MOTIVE} consistently outperforms strong self-verification and self-correction baselines. Further results show that reliable verification improves accept-or-rethink decisions and reduces unnecessary reasoning turns, enabling more reliable and efficient self-verification without an external judge.

Explore similar work

CardsList
  1. Improving Reasoning in Vision-Language Models via Perception Verified Self-Training

    Jun 20, 2026Sourabh Sharma, Sonam Gupta, SadbhawnaVLM ReasoningVLM Hallucination Mitigation

  2. SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning

    Jul 13, 2026Mingyuan Wu, Jingcheng Yang, Shengyi Qian +11Vision-Language ModelsMultimodal Reasoning

  3. See, Think, Learn: A Self-Taught Multimodal Reasoner

    Dec 2, 2025Sourabh Sharma, Sonam Gupta, SadbhawnaSelf-TrainingVLM Reasoning