On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a frozen or EMA version of themselves that receives privileged information. Its application to multimodal large language models (MLLMs), however, remains largely unexplored. Recent approaches use privileged visual information, such as image crops corresponding to a question, to improve fine-grained perception, but their gains are confined to tasks that benefit from such visual zooming and require either human-annotated grounding data or external teacher models. We introduce a different form of on-policy self-distillation for MLLMs that provides the teacher with textual, spatially grounded guidance identifying the visual elements relevant to a query. We use procedurally generated scenes with automatically available object identities and spatial coordinates, enabling scalable and annotation-free post-training. The teacher uses this spatial guidance to locate and integrate evidence from multiple relevant image regions, while the student learns to reproduce the resulting behavior from the image and question alone. Our approach consistently improves performance on counting, document and chart understanding benchmarks across multiple models. Importantly, although post-training uses only synthetic scenes, the resulting improvements transfer to real-world perception benchmarks, yielding a 3.23-point gain in average performance across CVBench, V*, ZoomBench, BLINK, HR-Bench, and MME-RealWorld. These results show that spatially grounded privileged information can induce broader perceptual capabilities through on-policy self-distillation, enabling substantial synthetic-to-real transfer beyond the task and data distribution used for post-training. Project page: https://github.com/sirkosophia/Where-OPD
Figures & tables
General
High-resolution
Model
CVB
MME-RW
BLINK
GQA
CountQA
ChartQA
DocVQA
OCRB
EvoChart
HallB
AMBER
V ∗
HR-4K
HR-8K
Zoom
Avg.
Qwen3.5-4B base model
87.45
64.11
66.28
69.32
32.00
79.32
96.00
87.30
68.32
69.40
88.80
82.72
86.25
80.00
50.65
73.86
Vision-OPD
87.53 + 0.08
75.26 + 11.15
62.02 − 4.26
68.23 − 1.09
21.27 − 10.73
78.52 − 0.80
95.98 − 0.02
87.00 − 0.30
67.28 − 1.04
65.83 − 3.57
86.32 − 2.48
91.62 + 8.90
83.50 − 2.75
80.25 + 0.25
65.21 + 14.56
74.39 + 0.53
OPD-V
85.97 − 1.48
75.76 + 11.65
61.91 − 4.37
67.61 − 1.71
26.11 − 5.89
81.12 + 1.80
96.00 ± 0.00
87.50 + 0.20
70.32 + 2.00
68.03 − 1.37
82.78 − 6.02
94.24 + 11.52
85.00 − 1.25
82.38 + 2.38
65.44 + 14.79
75.34 + 1.48
S 2 VOPD
87.91 + 0.46
75.59 + 11.48
64.70 − 1.58
68.61 − 0.71
31.87 − 0.13
79.88 + 0.56
95.98 − 0.02
87.50 + 0.20
69.44 + 1.12
65.72 − 3.68
88.20 − 0.60
88.48 + 5.76
85.12 − 1.13
83.88 + 3.88
56.09 + 5.44
75.26 + 1.40
Where-OPD
89.76 + 2.31
74.75 + 10.64
67.77 + 1.49
75.41 + 6.09
34.53 + 2.53
86.52 + 7.20
97.00 + 1.00
89.23 + 1.93
78.43 + 10.11
80.16 + 10.76
93.28 + 4.48
87.96 + 5.24
85.62 − 0.63
82.67 + 2.67
51.52 + 0.87
78.31 + 4.45
Table 1: Main results across three MLLMs. Accuracy (%) on 15 benchmarks; Avg. is their unweighted mean. The smaller numbers below each score show the change in accuracy relative to the corresponding base model. CVB: CVBench; MME-RW: MME-RealWorld; OCRB: OCRBench; HallB: HallusionBench; HR-4K/8K: HR-Bench 4K/8K; Zoom: ZoomBench. For Where-OPD on Qwen3.5-4B and Qwen3.5-9B as base model, each score is the mean of three training runs with different seeds; the standard deviation of Avg. across runs is ± 0.20 for both Qwen3.5-4B and Qwen3.5-9B.
Training signal
CVB
V ∗
HR-4K
HR-8K
Zoom
CountQA
EvoChart
Avg.
Qwen3.5-4B
87.45
82.72
86.25
80.00
50.65
32.00
68.32
69.63
SFT
87.68
89.01
83.38
79.12
50.18
31.81
66.72
69.70
GRPO
87.83
83.25
81.88
80.12
49.11
33.12
71.60
69.56
OPSD, no privileged information
87.53
82.72
85.00
81.00
50.41
32.46
68.32
69.63
OPSD, cropped-image hint
88.44
83.77
82.50
79.25
52.31
20.22
74.80
68.76
OPSD, answer-only hint
88.48
84.82
83.25
79.38
51.12
33.51
78.32
71.27
Table 2: Training signals and privileged hints on Qwen3.5-4B. SFT, GRPO, and the OPSD variants are trained on the same 3,000 synthetic scenes for the same number of steps as Where-OPD . Results are accuracy (%); Avg. is the unweighted mean across benchmarks present in the table. CVB: CVBench; HR-4K/8K: HR-Bench 4K/8K; Zoom: ZoomBench.
General
High-resolution
Model
CVB
MME-RW
BLINK
GQA
CountQA
ChartQA
DocVQA
OCRB
EvoChart
HallB
AMBER
V ∗
HR-4K
HR-8K
Zoom
Avg.
Qwen3.5-4B
89.16
70.30
71.23
69.32
46.86
86.28
96.45
87.30
77.36
69.82
88.80
82.72
88.12
81.50
50.65
77.06
Vision-OPD
88.51 − 0.65
75.26 + 4.96
67.75 − 3.48
68.23 − 1.09
42.74 − 4.12
86.04 − 0.24
96.24 − 0.21
87.00 − 0.30
75.52 − 1.84
67.09 − 2.73
86.32 − 2.48
92.15 + 9.43
83.50 − 4.62
81.00 − 0.50
65.21 + 14.56
77.50 + 0.45
Where-OPD
90.45 + 1.29
74.59 + 4.29
71.23 ± 0.00
75.39 + 6.07
54.12 + 7.26
86.56 + 0.28
97.01 + 0.56
89.20 + 1.90
78.64 + 1.28
81.60 + 11.78
94.37 + 5.57
88.48 + 5.76
86.62 − 1.50
83.25 + 1.75
53.61 + 2.96
80.34 + 3.28
Table 3: Best tested inference settings on Qwen3.5-4B. For each model and benchmark, the reported accuracy (%) is the best across thinking enabled or thinking disabled. Settings are selected separately for each cell; Avg. is the mean of the 15 selected scores. Smaller numbers show percentage-point changes from the base model. CVB: CVBench; MME-RW: MME-RealWorld; OCRB: OCRBench; HallB: HallusionBench; HR-4K/8K: HR-Bench 4K/8K; Zoom: ZoomBench.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Category
n
Qwen3.5-4B
Where-OPD
Δ
CVBench
Count
788
72.72
76.65
+3.93
Depth
600
95.50
96.33
+0.83
Distance
600
90.17
93.50
+3.33
Relation
650
95.38
96.77
+1.39
Overall
2,638
87.45
89.92
+2.47
Appendix
Table 4: Per-category accuracy on Qwen3.5-4B. Accuracy (%) for Qwen3.5-4B and Where-OPD across six benchmarks, broken down by their category labels. n is the number of examples in each category, and Δ is Where-OPD minus base accuracy. Bold indicates the higher score.
Benchmarks
CVB
V ∗
HR-4K
HR-8K
Zoom
CountQA
EvoChart
Avg.
Qwen3.5-4B base model
87.45
82.72
86.25
80.00
50.65
32.00
68.32
69.63
Training-scene resolution
10242 (ours)
89.92
88.48
85.50
83.25
51.01
33.77
78.64
72.94
7682
89.99
86.39
85.12
79.50
49.94
33.57
77.52
71.72
5122
88.78
85.86
84.12
81.12
50.41
34.69
78.64
71.95
Appendix
Table 5: Scene-generation ablations on Qwen3.5-4B. Top: the same scenes and questions rendered at different resolutions, with hint coordinates rescaled accordingly. Bottom: different ranges for the number of objects per scene. Results are accuracy (%); Avg. is the unweighted mean across seven benchmarks. The highlighted rows show the default setting in both blocks; bold marks the best result within each block. CVB: CVBench; HR-4K/8K: HR-Bench 4K/8K; Zoom: ZoomBench.
Benchmarks
Examples
Steps
CVB
V ∗
HR-4K
HR-8K
Zoom
CountQA
EvoChart
Avg.
Qwen3.5-4B
87.45
82.72
86.25
80.00
50.65
32.00
68.32
69.63
1,000
10
87.00
84.29
83.50
79.00
47.93
34.42
70.24
69.48
2,000
20
89.54
86.91
85.88
81.88
50.65
33.97
78.24
72.44
3,000
31
89.92
88.48
85.50
83.25
51.01
33.77
78.64
72.94
6,000
62
88.97
83.77
84.12
81.50
49.94
35.86
79.20
71.91
Appendix
Table 6: Training-set size. Accuracy (%) for Qwen3.5-4B (top) and Qwen3-VL-4B (bottom) after one epoch at each dataset size. The highlighted rows are the setting used in the main experiments. Avg. is the unweighted mean across seven benchmarks; bold marks the best dataset size within each backbone. CVB: CVBench; HR-4K/8K: HR-Bench 4K/8K; Zoom: ZoomBench.
Benchmarks
Teacher update rate η
CVB
V ∗
HR-4K
HR-8K
Zoom
CountQA
EvoChart
Avg.
Qwen3.5-4B
87.45
82.72
86.25
80.00
50.65
32.00
68.32
69.63
0 (frozen)
89.92
88.48
85.50
83.25
51.01
33.77
78.64
72.94
0.0001
89.61
89.01
83.12
81.50
51.36
35.47
77.92
72.57
0.001
89.31
86.91
84.75
80.38
51.01
34.75
77.84
72.14
0.01
89.27
84.82
84.25
81.62
51.01
34.75
76.88
71.80
Appendix
Table 7: Teacher update-rate study. Accuracy (%) for Qwen3.5-4B (top) and Qwen3-VL-4B (bottom) at different EMA update rates η ; η=0 leaves the teacher frozen. The highlighted rows are used in the main experiments. Avg. is the unweighted mean across seven benchmarks; bold marks the best update rate within each backbone. CVB: CVBench; HR-4K/8K: HR-Bench 4K/8K; Zoom: ZoomBench.
General
High-resolution
Model
CVB
MME-RW
BLINK
GQA
CountQA
ChartQA
DocVQA
OCRB
EvoChart
HallB
AMBER
V ∗
HR-4K
HR-8K
Zoom
Avg.
Qwen3.5-4B base model
89.16
70.30
71.23
64.86
46.86
86.28
96.45
86.90
77.36
69.82
88.61
82.72
88.12
81.50
46.15
76.42
Vision-OPD
88.51 − 0.65
70.43 + 0.13
67.75 − 3.48
62.26 − 2.60
42.74 − 4.12
86.04 − 0.24
96.24 − 0.21
85.80 − 1.10
75.52 − 1.84
67.09 − 2.73
83.88 − 4.73
92.15 + 9.43
83.50 − 4.62
81.00 − 0.50
57.99 + 11.84
76.06 − 0.36
OPD-V
86.35 − 2.81
69.81 − 0.49
66.12 − 5.11
58.45 − 6.41
45.03 − 1.83
84.04 − 2.24
96.09 − 0.36
84.70 − 2.20
74.40 − 2.96
64.25 − 5.57
78.13 − 10.48
89.53 + 6.81
84.25 − 3.87
79.75 − 1.75
57.28 + 11.13
74.55 − 1.88
S2VOPD
88.82 − 0.34
71.38 + 1.08
70.65 − 0.58
61.71 − 3.15
53.93 + 7.07
85.04 − 1.24
96.28 − 0.17
86.00 − 0.90
75.20 − 2.16
66.56 − 3.26
83.91 − 4.70
85.34 + 2.62
86.75 − 1.37
83.50 + 2.00
53.61 + 7.46
76.58 + 0.16
Where-OPD
90.45 + 1.29
71.50 + 1.20
71.23 ± 0.00
65.81 + 0.95
54.12 + 7.26
86.16 − 0.12
96.49 + 0.04
86.60 − 0.30
76.00 − 1.36
69.40 − 0.42
87.61 − 1.00
83.25 + 0.53
86.62 − 1.50
81.75 + 0.25
53.61 + 7.46
77.37 + 0.95
Appendix
Table 8: Evaluations under thinking enabled mode on Qwen3.5-4B. All models are evaluated with thinking enabled. Results are accuracy (%); Avg. is the unweighted mean across 15 benchmarks. The smaller numbers below each score show the percentage-point change from the base model evaluated in the same mode. CVB: CVBench; MME-RW: MME-RealWorld; OCRB: OCRBench; HallB: HallusionBench; HR-4K/8K: HR-Bench 4K/8K; Zoom: ZoomBench.
Model
L3
L7
L11
L15
L19
L23
L27
L31
Chosen
Qwen3.5-4B
0.68
0.27
0.64
0.94
0.92
0.76
1.15
0.54
L27
Vision-OPD
1.12
0.71
3.01
4.75
3.20
2.31
2.57
1.47
L15
Where-OPD
1.32
1.17
3.36
5.79
4.30
2.69
2.29
1.42
L15
Appendix
Table 9: Attention-layer selection on synthetic counting scenes. Mean enrichment of attention on the target objects at each full-attention layer, measured on a separate selection split of 158 held-out images with at least three targets. Enrichment of 1 means the targets receive attention proportional to their area. Chosen is the layer with the highest enrichment for each model. The base model is Qwen3.5-4B.
Chinese Information Processing Laboratory, Institute of Software, Chinese Academy of Sciences · University of Chinese Academy of Sciences · Xiaohongshu Inc.