cs.CVSep 29, 2026

On-Policy Visual Evidence Distillation

Authors: Shaohang Wei, Feifan Song, Guangyue Peng, Wenhao Yu, Wei Li, Wen Luo, Yang Xu, Yufan Shen, +4 more

Organizations: Peking University · CUHK · Nanjing University · Tencent

Abstract

Visual agents solve problems by interleaving reasoning with image operations, and on-policy distillation (OPD) provides guidance from a strong teacher on student-generated interaction trajectories. However, image operations change the evidence available for subsequent reasoning, so local errors in evidence acquisition (Acquire), reading (Read), or answer grounding (Ground) can propagate through the trajectory and lead to incorrect answers. Existing multimodal OPD methods primarily construct or contrast auxiliary views of the original image to strengthen supervision, without explicitly modeling the connections between student actions, resulting observations, and subsequent reasoning. This limits their ability to provide corrections tailored to different failure stages. We introduce Reflection on Visual Evidence (ReVuE), an on-policy distillation method for visual agents. ReVuE compares multiple student-generated trajectories for the same query, summarizes the observed visual evidence, and diagnoses the first failure across the Acquire, Read, and Ground stages. The resulting reflections provide training-time context for the teacher. We group and reweight token-level distillation losses according to how strongly these reflections affect the teacher's predictions. This design translates trajectory-level evidence diagnosis into targeted token-level supervision, guiding students to improve their visual evidence acquisition and reasoning. Across 11 benchmarks spanning the Qwen2.5-VL and InternVL3.5 model families, ReVuE outperforms all evaluated OPD baselines in weighted-average scores for perception, mathematical reasoning, and general tasks. ReVuE also reduces redundancy in reasoning and tool calls while improving tool-call accuracy and task accuracy. Code is available at https://github.com/sylvain-wei/ReVuE

Figures & tables

Appendix figures & tables22 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ViCuR: Visual Cues as Recoverable Privilege for Multimodal On-Policy Distillation

    Jun 4, 2026Kanghui Tian, Siyuan Liu, Ziang Yan +3Multimodal ReasoningVisual Cues

  2. OPD-Aha: From Linguistic Momentum to Visual Reflection in Multimodal On-Policy Distillation

    Sep 15, 2026Chenhao Qiu, Dawei Li, Yechao Zhang +2Multimodal ReasoningTeacher

  3. VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation

    Jul 30, 2026Kangning Zhang, Yixing Li, Shuai Shao +9Dataset DistillationVisual Evidence