Visual agents solve problems by interleaving reasoning with image operations, and on-policy distillation (OPD) provides guidance from a strong teacher on student-generated interaction trajectories. However, image operations change the evidence available for subsequent reasoning, so local errors in evidence acquisition (Acquire), reading (Read), or answer grounding (Ground) can propagate through the trajectory and lead to incorrect answers. Existing multimodal OPD methods primarily construct or contrast auxiliary views of the original image to strengthen supervision, without explicitly modeling the connections between student actions, resulting observations, and subsequent reasoning. This limits their ability to provide corrections tailored to different failure stages. We introduce Reflection on Visual Evidence (ReVuE), an on-policy distillation method for visual agents. ReVuE compares multiple student-generated trajectories for the same query, summarizes the observed visual evidence, and diagnoses the first failure across the Acquire, Read, and Ground stages. The resulting reflections provide training-time context for the teacher. We group and reweight token-level distillation losses according to how strongly these reflections affect the teacher's predictions. This design translates trajectory-level evidence diagnosis into targeted token-level supervision, guiding students to improve their visual evidence acquisition and reasoning. Across 11 benchmarks spanning the Qwen2.5-VL and InternVL3.5 model families, ReVuE outperforms all evaluated OPD baselines in weighted-average scores for perception, mathematical reasoning, and general tasks. ReVuE also reduces redundancy in reasoning and tool calls while improving tool-call accuracy and task accuracy. Code is available at https://github.com/sylvain-wei/ReVuE
Figures & tables
Figure 1: ReVuE overview (with HRBench-8K results) and an agentic vision example.
Figure 2: Sampling and training pipeline of ReVuE .
Figure 3: Token impact on a trajectory with a reading error. The student misreads two islands as one, answering Saint Lucia instead of Saint Kitts and Nevis . Reflection identifies two islands and lowers support for the underlined island , one , and first Lucia . Blue / orange denote increased/decreased teacher support; darker shading indicates larger absolute log-probability shifts. Both evaluations score the same student trajectory; the final Lucia changes little under the fixed erroneous prefix.
Figure 5: Response length, accuracy, and tool use for Qwen2.5-VL-7B.
Model / Method
Perception
Math
General
HR-4K
HR-8K
V *
Tree Bench
Visual Probe
Wtd. Avg.
Math Vista
Math Verse
Visu Logic
Wtd. Avg.
Hallu Bench
CQA Pro
Info VQA
Wtd. Avg.
Off-the-Shelf Models
GPT-4o
61.00
54.00
61.78
49.88
23.88
50.28
58.83
39.21
25.20
41.22
51.37
28.67
71.17
53.28
Gemini3.1FL
46.00
43.00
64.92
51.85
27.96
43.89
80.80
77.92
32.40
62.63
59.92
37.03
83.50
63.57
Qwen2.5-32B
75.13
69.25
78.01
48.40
45.05
63.89
77.00
53.05
25.90
51.90
50.74
30.02
83.10
59.29
Qwen3-30B-T
77.13
71.38
80.10
45.43
36.89
63.26
80.20
66.12
25.80
56.71
61.82
35.46
85.68
64.45
Table 1: Main results. Bold marks the best OPD result per column within each model family, including ties. Wtd. Avg. uses benchmark sample counts as weights. Baseline details and model labels are provided in Appendix B.1 – B.2 . Evaluation benchmark details are provided in Appendix F .
Figure 4: Token-impact sparsity
Method
Perception ↑
Math ↑
General ↑
Vanilla OPD
62.61
47.67
53.28
+reflection
63.70
47.78
54.42
+impact reweight
65.01
48.49
54.14
Table 2: Component ablation on Qwen2.5-VL-7B.
Method
Perception ↑
Math ↑
General ↑
Vanilla OPD
62.61
47.67
53.28
Random select 20%
63.64
47.42
54.10
Mask random 20%
63.85
47.35
53.78
Mask low 20%
65.18
48.42
53.97
Mask top 20%
62.26
47.17
52.51
ReVuE (Ours)
65.01
48.49
54.14
Table 3: Token-impact interventions.
Fraction
Perception ↑
Math ↑
General ↑
Top 5%
64.70
47.49
54.12
Top 10%
64.55
46.92
53.85
Top 20% ( ReVuE )
65.01
48.49
54.14
Top 50%
64.33
47.42
54.18
Top 100%
63.70
47.78
54.42
Table 4: Sensitivity to α .
Comparator
n
Agreement (%)
κ
Human annotator
96
96.88
0.957
Qwen (rerun)
744
99.2
0.987
Gemini 3.8 Flash
736
93.2
0.892
GPT-5.6 Luna
736
92.8
0.886
Table 5: Critic stage consistency.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: A concrete visual-evidence reflection supplied to the teacher. The mixed group contains seven correct attempts and one incorrect attempt; one of each is shown with its post-tool reasoning and answer. Both attempts receive identical tool crops. The Anchor uses the incorrect attempt’s own crop, and the Break Point records the critic’s Ground diagnosis.
Component
Parameter
Value
OPD objective
Loss
Reverse KL
Vocabulary support
Teacher’s top-32 candidates
OPD rollouts
Temperature for code tokens
0.0
Temperature for other tokens
1.0
Top- p
0.9
Top- k
50
Appendix
Table 6: Training and evaluation settings.
Human \ Critic
Correct
Acquire
Read
Ground
Other
Correct
32
0
0
0
0
Acquire
0
28
0
0
0
Read
0
3
23
0
0
Ground
0
0
0
9
0
Other
0
0
0
0
1
Appendix
Table 7: Final critic stage agreement.
Figure 7: Human annotation interface demo: the upper view presents the sample identifiers, original image, question, and reference answer. The page also provides progress tracking, sample filtering, navigation, and JSONL import and export controls.
Figure 8: Human annotation interface demo: the lower view of the same sample presents the rollout, the initially collapsed judge label, six annotation options, and a notes field. The options comprise five diagnostic labels and a SKIP option for cases that cannot be judged reliably.
Method
HR-4K
HR-8K
V *
Tree Bench
Visual Probe
Wtd. Avg.
Vanilla OPD
75.40
70.50
81.20
38.02
42.91
62.61
+ Visual-evidence reflection
76.25
72.50
82.20
37.78
44.08
63.70
+ Impact reweighting ( ReVuE )
77.10
74.00
82.20
41.12
44.66
65.01
Appendix
Table 8: Complete component ablation on Qwen2.5-VL-7B. Rows cumulatively add visual-evidence reflection and impact reweighting, and Wtd. Avg. uses the benchmark-sample weighting of Table 1 . MathVista and MathVerse denote the Mini and vision-only splits, respectively. Higher is better for every metric; gray denotes Vanilla OPD, blue marks ReVuE , and bold indicates column-best results, including ties.
Method
HR-4K
HR-8K
V *
Tree Bench
Visual Probe
Wtd. Avg.
Vanilla OPD
75.40
70.50
81.20
38.02
42.91
62.61
Random select 20%
75.40
72.62
80.10
38.02
45.44
63.64
Mask random 20%
76.25
73.75
81.20
37.28
43.69
63.85
Mask low 20%
77.00
75.00
82.20
39.51
45.44
65.18
Mask top 20%
75.00
71.62
78.01
37.78
41.36
62.26
ReVuE (Ours)
77.10
74.00
82.20
41.12
44.66
65.01
Appendix
Table 9: Complete token-impact intervention results on Qwen2.5-VL-7B. Wtd. Avg. uses the benchmark-sample weighting of Table 1 . MathVista and MathVerse denote the Mini and vision-only splits, respectively. All scores are percentages and higher is better; gray denotes Vanilla OPD, blue marks ReVuE , and bold indicates column-best results, including ties.
Method
Masked positions
Retained supervision
Random select 20%
None
All positions; random 20% and remaining 80% receive group loss weights 0.5:0.5
Mask random 20%
Random 20%
Remaining 80% receive total group loss weight 0.5 , shared uniformly
Mask low 20%
Lowest-impact 20%
Highest-impact 80%, including the entire High group
Mask top 20%
Highest-impact 20%
Lowest-impact 80%; the entire High group is masked
ReVuE (Ours)
None
All positions; highest-impact 20% and remaining 80% receive group loss weights 0.5:0.5
Appendix
Table 10: Variant definitions for token-impact interventions. Percentages refer to supervised positions, and masking removes the selected positions from the distillation loss. Random select replaces the impact-selected high-weight group with a random group while retaining supervision at every position.
High-impact fraction
HR-4K
HR-8K
V *
Tree Bench
Visual Probe
Wtd. Avg.
Top 5%
77.00
73.12
83.25
40.49
44.66
64.70
Top 10%
75.88
73.75
82.72
40.99
44.47
64.55
Top 20% ( ReVuE )
77.10
74.00
82.20
41.12
44.66
65.01
Top 50%
76.50
73.50
80.63
39.26
44.85
64.33
Top 100%
76.25
72.50
82.20
37.78
44.08
63.70
Appendix
Table 11: Sensitivity to the high-impact token fraction on Qwen2.5-VL-7B. Wtd. Avg. is weighted by benchmark sample counts. Sample counts are (800, 800, 191, 405, 515), (1000, 788, 1000), and (1129, 1948, 2801) for perception, math, and general benchmarks, respectively. MathVista and MathVerse denote the Mini and vision-only splits, respectively. All scores are percentages and higher is better; blue marks the 20% setting of ReVuE , and bold indicates column-best results, including ties.
Figure 9: Response length, accuracy, and tool use on HRBench 4K for Qwen2.5-VL-7B.
Figure 10: Response length, accuracy, and tool use on HRBench 8K for Qwen2.5-VL-7B.
Figure 11: Response length, accuracy, and tool use on V* Bench for Qwen2.5-VL-7B.
Figure 12: Response length, accuracy, and tool use on TreeBench for Qwen2.5-VL-7B.
Figure 13: Response length, accuracy, and tool use on VisualProbe for Qwen2.5-VL-7B.
Reflection status
Rollout slots
Share (%)
Complete reflection
134,784
78.36
Break-Point-only reflection
26,848
15.61
No reflection (base teacher)
10,368
6.03
Any reflection (subtotal)
161,632
93.97
Appendix
Table 12: Reflection coverage in the main training run. All percentages use 172,000 rollout slots as the denominator; the final row sums the first two rows.
Context
Correct / 311
Accuracy (%)
Δ Accuracy (%)
Original predictions
172
55.31
—
Anchor only
187
60.13
+4.82
Anchor + Break Point
288
92.60
+37.30
Appendix
Table 13: Offline reflection comparison on fixed paired cases. Original predictions are the predictions before adding reflection, and accuracy changes are computed from unrounded values.
Figure 14: Per-position gradient weights at λ=0.5 . Blue and gray curves show the High and Low gradient multipliers relative to uniform token averaging; the black dotted line marks the unit baseline. The continuous curves ignore group-size rounding and cover 0.1≤α≤0.9 . Markers at α=0.2 give the exact multipliers 2.5 and 0.625 for the 100-position example.
Figure 15: In Acquire , the crop misses the queried dogs; in Read , the model describes a curve above the x-axis as lying below it; in Ground , it reads the table correctly but miscalculates the sum. The panels retain post-crop thinking in (a), full thinking in (b), and the complete reasoning-bearing final answer in (c). Red text marks errors and their consequences.
Figure 16: Medium Blue is the fourth-largest slice, making it the lower median of six slices. The left trajectory misreads the rank ( Read ); the right trajectory reads it correctly but misapplies the lower-median rule ( Ground ). Both complete thinking traces and final answers are shown, with errors in red.
Figure 17: Both trajectories answer B (RRATs), whereas the reference answer is A (AXPs and SGRs). The left misidentifies green lines as age contours; the right correctly places AXPs and SGRs above the age contour but interprets that position as older. Their different crops and all natural-language thinking are shown; tool code is summarized.
Figure 18: Token impact on visual target selection. The student crops the foreground person and reports light purple or lavender instead of red . Reflection identifies the background person as the target and lowers support for the underlined foreground and left in the crop plan. Blue / orange denote increased/decreased teacher support; darker shading indicates larger absolute log-probability shifts on the same student trajectory.
Figure 19: Token impact on age inference. The student correctly places AXPs and SGRs above the τc=105 yr contour but infers older ages, incorrectly answering RRATs . Reflection links this position to younger ages, increasing support for the underlined correct relation above and decreasing support for the mistaken comparison exceed . Blue / orange denote increased/decreased teacher support; darker shading indicates larger absolute log-probability shifts on the same student trajectory.
On-policy distillation (OPD) improves reasoning by training a student on trajectories sampled from its own policy under supervision from a teacher. In multimodal reasoning, a common extension is to use a privileged teacher that observes training-time-only signals such as reference answers or rationales. However, such answer-side privilege creates a train-test mismatch: the teacher's supervision may depend on signals unavailable to the student, encouraging shortcut imitation rather than visually grounded reasoning. We propose ViCuR, a visually grounded privileged-teacher distillation framework that replaces answer-side privilege with visual cues (query-related evidence in the input). Because these cues are derived from the same visual input available at inference, their evidence is recoverable by the student. To support this, ViCuR introduces a lightweight cue recovery module that uses dedicated sink-token cross-attention during prefill to aggregate task-relevant visual evidence into an internal representation, without changing the inference interface or requiring auxiliary cue-generation losses. Across seven benchmarks with Qwen3-VL-2B and 8B students, ViCuR consistently improves over answer-based on-policy self-distillation by +1.19 and +1.24 on overall average performance. It also extends naturally to stronger-teacher OPD, surpassing OPD baselines by +0.64 and +1.08, with consistent out-of-domain gains at the 8B scale. These results show that, in multimodal on-policy distillation, the design of teacher privilege is as important as teacher strength.
Kanghui Tian, Siyuan Liu, Ziang Yan +3
1Shanghai AI Laboratory · 3Nanjing University · 2Fudan University
Privileged on-policy distillation improves multimodal reasoning by allowing a teacher to evaluate student trajectories using rich, training-only visual evidence. Both models score these trajectories while conditioning on the same student-generated prefix. When a student misinterprets an image early in a response, this accumulating erroneous rationale eventually pulls the teacher away from its visual evidence. The teacher and student converge on the same hallucination, causing standard cross-model supervision to collapse precisely where correction is most needed. We find that the teacher's visual corrective preference is not lost under this misleading agreement. Comparing the predictions of the identical teacher given the real image and a visual null reveals that the privileged evidence still pushes the model toward the correct interpretation. We introduce OPD-Aha, which reconstructs the distillation target directly from this isolated visual preference rather than relying on the fragile teacher-student discrepancy. This reconstructed target aggressively suppresses continuations that contradict the image. Trained with this objective, students learn to naturally interrupt their own flawed reasoning with reflection tokens such as wait and actually. After reflection, subsequent generation relies less on the accumulated erroneous text and more on the visual evidence. Correcting these trajectories mid-generation fundamentally alters the reasoning process, yielding broad and consistent improvements across diverse fine-grained perception and complex multimodal reasoning benchmarks. Our code and models are available at https://github.com/Echochef/OPD-Aha.
Chenhao Qiu, Dawei Li, Yechao Zhang +2
Arizona State University · University of Virginia · Stevens Institute of Technology
Multimodal on-policy distillation (OPD) transfers fine-grained visual knowledge by supervising student-generated trajectories with a privileged-view teacher. Yet its next-token corrections are source-mixed, combining visual signals with linguistic priors and teacher-specific effects. The key challenge is to estimate which corrections are supported by visual evidence, not merely where or how strongly to distill. We introduce Visual Attribution Distillation (VAD), a counterfactual target-reconstruction algorithm that estimates the visually attributable part of a teacher correction. At each student-generated prefix, VAD evaluates the same fixed teacher with the relevant evidence present and removed. The corresponding change in centered log-probabilities defines ut, a signed proxy for the visual evidence direction that estimates how revealing the evidence supports or refutes candidate tokens. VAD projects the original correction onto this proxy to obtain an intervention-aligned component and a proxy-unexplained residual, then reconstructs a student-anchored target from the former. During training, this reconstructed target supplies the primary supervision signal, while the privileged teacher contributes a weak regularizer. Across six fine-grained visual benchmarks at 4B and 9B scales, VAD outperforms direct privileged-view distillation and visual-advantage weighting. Token- level and controlled-target analyses show that the proxy-aligned component is enriched in task-relevant visual corrections and yields stronger target shifts, especially when evidence refutes a mistaken answer. These results support counterfactual target reconstruction as an effective alternative to source-mixed supervision.
Kangning Zhang, Yixing Li, Shuai Shao +9
Shanghai Jiao Tong University · Xiaohongshu Inc. · The Chinese University of Hong Kong +1