Visual agents solve problems by interleaving reasoning with image operations, and on-policy distillation (OPD) provides guidance from a strong teacher on student-generated interaction trajectories. However, image operations change the evidence available for subsequent reasoning, so local errors in evidence acquisition (Acquire), reading (Read), or answer grounding (Ground) can propagate through the trajectory and lead to incorrect answers. Existing multimodal OPD methods primarily construct or contrast auxiliary views of the original image to strengthen supervision, without explicitly modeling the connections between student actions, resulting observations, and subsequent reasoning. This limits their ability to provide corrections tailored to different failure stages. We introduce Reflection on Visual Evidence (ReVuE), an on-policy distillation method for visual agents. ReVuE compares multiple student-generated trajectories for the same query, summarizes the observed visual evidence, and diagnoses the first failure across the Acquire, Read, and Ground stages. The resulting reflections provide training-time context for the teacher. We group and reweight token-level distillation losses according to how strongly these reflections affect the teacher's predictions. This design translates trajectory-level evidence diagnosis into targeted token-level supervision, guiding students to improve their visual evidence acquisition and reasoning. Across 11 benchmarks spanning the Qwen2.5-VL and InternVL3.5 model families, ReVuE outperforms all evaluated OPD baselines in weighted-average scores for perception, mathematical reasoning, and general tasks. ReVuE also reduces redundancy in reasoning and tool calls while improving tool-call accuracy and task accuracy. Code is available at https://github.com/sylvain-wei/ReVuE
Figures & tables
Figure 1: ReVuE overview (with HRBench-8K results) and an agentic vision example.
Figure 2: Sampling and training pipeline of ReVuE .
Figure 3: Token impact on a trajectory with a reading error. The student misreads two islands as one, answering Saint Lucia instead of Saint Kitts and Nevis . Reflection identifies two islands and lowers support for the underlined island , one , and first Lucia . Blue / orange denote increased/decreased teacher support; darker shading indicates larger absolute log-probability shifts. Both evaluations score the same student trajectory; the final Lucia changes little under the fixed erroneous prefix.
Figure 5: Response length, accuracy, and tool use for Qwen2.5-VL-7B.
Model / Method
Perception
Math
General
HR-4K
HR-8K
V *
Tree Bench
Visual Probe
Wtd. Avg.
Math Vista
Math Verse
Visu Logic
Wtd. Avg.
Hallu Bench
CQA Pro
Info VQA
Wtd. Avg.
Off-the-Shelf Models
GPT-4o
61.00
54.00
61.78
49.88
23.88
50.28
58.83
39.21
25.20
41.22
51.37
28.67
71.17
53.28
Gemini3.1FL
46.00
43.00
64.92
51.85
27.96
43.89
80.80
77.92
32.40
62.63
59.92
37.03
83.50
63.57
Qwen2.5-32B
75.13
69.25
78.01
48.40
45.05
63.89
77.00
53.05
25.90
51.90
50.74
30.02
83.10
59.29
Qwen3-30B-T
77.13
71.38
80.10
45.43
36.89
63.26
80.20
66.12
25.80
56.71
61.82
35.46
85.68
64.45
Table 1: Main results. Bold marks the best OPD result per column within each model family, including ties. Wtd. Avg. uses benchmark sample counts as weights. Baseline details and model labels are provided in Appendix B.1 – B.2 . Evaluation benchmark details are provided in Appendix F .
Figure 4: Token-impact sparsity
Method
Perception ↑
Math ↑
General ↑
Vanilla OPD
62.61
47.67
53.28
+reflection
63.70
47.78
54.42
+impact reweight
65.01
48.49
54.14
Table 2: Component ablation on Qwen2.5-VL-7B.
Method
Perception ↑
Math ↑
General ↑
Vanilla OPD
62.61
47.67
53.28
Random select 20%
63.64
47.42
54.10
Mask random 20%
63.85
47.35
53.78
Mask low 20%
65.18
48.42
53.97
Mask top 20%
62.26
47.17
52.51
ReVuE (Ours)
65.01
48.49
54.14
Table 3: Token-impact interventions.
Fraction
Perception ↑
Math ↑
General ↑
Top 5%
64.70
47.49
54.12
Top 10%
64.55
46.92
53.85
Top 20% ( ReVuE )
65.01
48.49
54.14
Top 50%
64.33
47.42
54.18
Top 100%
63.70
47.78
54.42
Table 4: Sensitivity to α .
Comparator
n
Agreement (%)
κ
Human annotator
96
96.88
0.957
Qwen (rerun)
744
99.2
0.987
Gemini 3.8 Flash
736
93.2
0.892
GPT-5.6 Luna
736
92.8
0.886
Table 5: Critic stage consistency.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: A concrete visual-evidence reflection supplied to the teacher. The mixed group contains seven correct attempts and one incorrect attempt; one of each is shown with its post-tool reasoning and answer. Both attempts receive identical tool crops. The Anchor uses the incorrect attempt’s own crop, and the Break Point records the critic’s Ground diagnosis.
Component
Parameter
Value
OPD objective
Loss
Reverse KL
Vocabulary support
Teacher’s top-32 candidates
OPD rollouts
Temperature for code tokens
0.0
Temperature for other tokens
1.0
Top- p
0.9
Top- k
50
Appendix
Table 6: Training and evaluation settings.
Human \ Critic
Correct
Acquire
Read
Ground
Other
Correct
32
0
0
0
0
Acquire
0
28
0
0
0
Read
0
3
23
0
0
Ground
0
0
0
9
0
Other
0
0
0
0
1
Appendix
Table 7: Final critic stage agreement.
Figure 7: Human annotation interface demo: the upper view presents the sample identifiers, original image, question, and reference answer. The page also provides progress tracking, sample filtering, navigation, and JSONL import and export controls.
Figure 8: Human annotation interface demo: the lower view of the same sample presents the rollout, the initially collapsed judge label, six annotation options, and a notes field. The options comprise five diagnostic labels and a SKIP option for cases that cannot be judged reliably.
Method
HR-4K
HR-8K
V *
Tree Bench
Visual Probe
Wtd. Avg.
Vanilla OPD
75.40
70.50
81.20
38.02
42.91
62.61
+ Visual-evidence reflection
76.25
72.50
82.20
37.78
44.08
63.70
+ Impact reweighting ( ReVuE )
77.10
74.00
82.20
41.12
44.66
65.01
Appendix
Table 8: Complete component ablation on Qwen2.5-VL-7B. Rows cumulatively add visual-evidence reflection and impact reweighting, and Wtd. Avg. uses the benchmark-sample weighting of Table 1 . MathVista and MathVerse denote the Mini and vision-only splits, respectively. Higher is better for every metric; gray denotes Vanilla OPD, blue marks ReVuE , and bold indicates column-best results, including ties.
Method
HR-4K
HR-8K
V *
Tree Bench
Visual Probe
Wtd. Avg.
Vanilla OPD
75.40
70.50
81.20
38.02
42.91
62.61
Random select 20%
75.40
72.62
80.10
38.02
45.44
63.64
Mask random 20%
76.25
73.75
81.20
37.28
43.69
63.85
Mask low 20%
77.00
75.00
82.20
39.51
45.44
65.18
Mask top 20%
75.00
71.62
78.01
37.78
41.36
62.26
ReVuE (Ours)
77.10
74.00
82.20
41.12
44.66
65.01
Appendix
Table 9: Complete token-impact intervention results on Qwen2.5-VL-7B. Wtd. Avg. uses the benchmark-sample weighting of Table 1 . MathVista and MathVerse denote the Mini and vision-only splits, respectively. All scores are percentages and higher is better; gray denotes Vanilla OPD, blue marks ReVuE , and bold indicates column-best results, including ties.
Method
Masked positions
Retained supervision
Random select 20%
None
All positions; random 20% and remaining 80% receive group loss weights 0.5:0.5
Mask random 20%
Random 20%
Remaining 80% receive total group loss weight 0.5 , shared uniformly
Mask low 20%
Lowest-impact 20%
Highest-impact 80%, including the entire High group
Mask top 20%
Highest-impact 20%
Lowest-impact 80%; the entire High group is masked
ReVuE (Ours)
None
All positions; highest-impact 20% and remaining 80% receive group loss weights 0.5:0.5
Appendix
Table 10: Variant definitions for token-impact interventions. Percentages refer to supervised positions, and masking removes the selected positions from the distillation loss. Random select replaces the impact-selected high-weight group with a random group while retaining supervision at every position.
High-impact fraction
HR-4K
HR-8K
V *
Tree Bench
Visual Probe
Wtd. Avg.
Top 5%
77.00
73.12
83.25
40.49
44.66
64.70
Top 10%
75.88
73.75
82.72
40.99
44.47
64.55
Top 20% ( ReVuE )
77.10
74.00
82.20
41.12
44.66
65.01
Top 50%
76.50
73.50
80.63
39.26
44.85
64.33
Top 100%
76.25
72.50
82.20
37.78
44.08
63.70
Appendix
Table 11: Sensitivity to the high-impact token fraction on Qwen2.5-VL-7B. Wtd. Avg. is weighted by benchmark sample counts. Sample counts are (800, 800, 191, 405, 515), (1000, 788, 1000), and (1129, 1948, 2801) for perception, math, and general benchmarks, respectively. MathVista and MathVerse denote the Mini and vision-only splits, respectively. All scores are percentages and higher is better; blue marks the 20% setting of ReVuE , and bold indicates column-best results, including ties.
Figure 9: Response length, accuracy, and tool use on HRBench 4K for Qwen2.5-VL-7B.
Figure 10: Response length, accuracy, and tool use on HRBench 8K for Qwen2.5-VL-7B.
Figure 11: Response length, accuracy, and tool use on V* Bench for Qwen2.5-VL-7B.
Figure 12: Response length, accuracy, and tool use on TreeBench for Qwen2.5-VL-7B.
Figure 13: Response length, accuracy, and tool use on VisualProbe for Qwen2.5-VL-7B.
Reflection status
Rollout slots
Share (%)
Complete reflection
134,784
78.36
Break-Point-only reflection
26,848
15.61
No reflection (base teacher)
10,368
6.03
Any reflection (subtotal)
161,632
93.97
Appendix
Table 12: Reflection coverage in the main training run. All percentages use 172,000 rollout slots as the denominator; the final row sums the first two rows.
Context
Correct / 311
Accuracy (%)
Δ Accuracy (%)
Original predictions
172
55.31
—
Anchor only
187
60.13
+4.82
Anchor + Break Point
288
92.60
+37.30
Appendix
Table 13: Offline reflection comparison on fixed paired cases. Original predictions are the predictions before adding reflection, and accuracy changes are computed from unrounded values.
Figure 14: Per-position gradient weights at λ=0.5 . Blue and gray curves show the High and Low gradient multipliers relative to uniform token averaging; the black dotted line marks the unit baseline. The continuous curves ignore group-size rounding and cover 0.1≤α≤0.9 . Markers at α=0.2 give the exact multipliers 2.5 and 0.625 for the 100-position example.
Figure 15: In Acquire , the crop misses the queried dogs; in Read , the model describes a curve above the x-axis as lying below it; in Ground , it reads the table correctly but miscalculates the sum. The panels retain post-crop thinking in (a), full thinking in (b), and the complete reasoning-bearing final answer in (c). Red text marks errors and their consequences.
Figure 16: Medium Blue is the fourth-largest slice, making it the lower median of six slices. The left trajectory misreads the rank ( Read ); the right trajectory reads it correctly but misapplies the lower-median rule ( Ground ). Both complete thinking traces and final answers are shown, with errors in red.
Figure 17: Both trajectories answer B (RRATs), whereas the reference answer is A (AXPs and SGRs). The left misidentifies green lines as age contours; the right correctly places AXPs and SGRs above the age contour but interprets that position as older. Their different crops and all natural-language thinking are shown; tool code is summarized.
Figure 18: Token impact on visual target selection. The student crops the foreground person and reports light purple or lavender instead of red . Reflection identifies the background person as the target and lowers support for the underlined foreground and left in the crop plan. Blue / orange denote increased/decreased teacher support; darker shading indicates larger absolute log-probability shifts on the same student trajectory.
Figure 19: Token impact on age inference. The student correctly places AXPs and SGRs above the τc=105 yr contour but infers older ages, incorrectly answering RRATs . Reflection links this position to younger ages, increasing support for the underlined correct relation above and decreasing support for the mistaken comparison exceed . Blue / orange denote increased/decreased teacher support; darker shading indicates larger absolute log-probability shifts on the same student trajectory.