On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a frozen or EMA version of themselves that receives privileged information. Its application to multimodal large language models (MLLMs), however, remains largely unexplored. Recent approaches use privileged visual information, such as image crops corresponding to a question, to improve fine-grained perception, but their gains are confined to tasks that benefit from such visual zooming and require either human-annotated grounding data or external teacher models. We introduce a different form of on-policy self-distillation for MLLMs that provides the teacher with textual, spatially grounded guidance identifying the visual elements relevant to a query. We use procedurally generated scenes with automatically available object identities and spatial coordinates, enabling scalable and annotation-free post-training. The teacher uses this spatial guidance to locate and integrate evidence from multiple relevant image regions, while the student learns to reproduce the resulting behavior from the image and question alone. Our approach consistently improves performance on counting, document and chart understanding benchmarks across multiple models. Importantly, although post-training uses only synthetic scenes, the resulting improvements transfer to real-world perception benchmarks, yielding a 3.23-point gain in average performance across CVBench, V*, ZoomBench, BLINK, HR-Bench, and MME-RealWorld. These results show that spatially grounded privileged information can induce broader perceptual capabilities through on-policy self-distillation, enabling substantial synthetic-to-real transfer beyond the task and data distribution used for post-training. Project page: https://github.com/sirkosophia/Where-OPD
Figures & tables
General
High-resolution
Model
CVB
MME-RW
BLINK
GQA
CountQA
ChartQA
DocVQA
OCRB
EvoChart
HallB
AMBER
V ∗
HR-4K
HR-8K
Zoom
Avg.
Qwen3.5-4B base model
87.45
64.11
66.28
69.32
32.00
79.32
96.00
87.30
68.32
69.40
88.80
82.72
86.25
80.00
50.65
73.86
Vision-OPD
87.53 + 0.08
75.26 + 11.15
62.02 − 4.26
68.23 − 1.09
21.27 − 10.73
78.52 − 0.80
95.98 − 0.02
87.00 − 0.30
67.28 − 1.04
65.83 − 3.57
86.32 − 2.48
91.62 + 8.90
83.50 − 2.75
80.25 + 0.25
65.21 + 14.56
74.39 + 0.53
OPD-V
85.97 − 1.48
75.76 + 11.65
61.91 − 4.37
67.61 − 1.71
26.11 − 5.89
81.12 + 1.80
96.00 ± 0.00
87.50 + 0.20
70.32 + 2.00
68.03 − 1.37
82.78 − 6.02
94.24 + 11.52
85.00 − 1.25
82.38 + 2.38
65.44 + 14.79
75.34 + 1.48
S 2 VOPD
87.91 + 0.46
75.59 + 11.48
64.70 − 1.58
68.61 − 0.71
31.87 − 0.13
79.88 + 0.56
95.98 − 0.02
87.50 + 0.20
69.44 + 1.12
65.72 − 3.68
88.20 − 0.60
88.48 + 5.76
85.12 − 1.13
83.88 + 3.88
56.09 + 5.44
75.26 + 1.40
Where-OPD
89.76 + 2.31
74.75 + 10.64
67.77 + 1.49
75.41 + 6.09
34.53 + 2.53
86.52 + 7.20
97.00 + 1.00
89.23 + 1.93
78.43 + 10.11
80.16 + 10.76
93.28 + 4.48
87.96 + 5.24
85.62 − 0.63
82.67 + 2.67
51.52 + 0.87
78.31 + 4.45
Table 1: Main results across three MLLMs. Accuracy (%) on 15 benchmarks; Avg. is their unweighted mean. The smaller numbers below each score show the change in accuracy relative to the corresponding base model. CVB: CVBench; MME-RW: MME-RealWorld; OCRB: OCRBench; HallB: HallusionBench; HR-4K/8K: HR-Bench 4K/8K; Zoom: ZoomBench. For Where-OPD on Qwen3.5-4B and Qwen3.5-9B as base model, each score is the mean of three training runs with different seeds; the standard deviation of Avg. across runs is ± 0.20 for both Qwen3.5-4B and Qwen3.5-9B.
Training signal
CVB
V ∗
HR-4K
HR-8K
Zoom
CountQA
EvoChart
Avg.
Qwen3.5-4B
87.45
82.72
86.25
80.00
50.65
32.00
68.32
69.63
SFT
87.68
89.01
83.38
79.12
50.18
31.81
66.72
69.70
GRPO
87.83
83.25
81.88
80.12
49.11
33.12
71.60
69.56
OPSD, no privileged information
87.53
82.72
85.00
81.00
50.41
32.46
68.32
69.63
OPSD, cropped-image hint
88.44
83.77
82.50
79.25
52.31
20.22
74.80
68.76
OPSD, answer-only hint
88.48
84.82
83.25
79.38
51.12
33.51
78.32
71.27
Table 2: Training signals and privileged hints on Qwen3.5-4B. SFT, GRPO, and the OPSD variants are trained on the same 3,000 synthetic scenes for the same number of steps as Where-OPD . Results are accuracy (%); Avg. is the unweighted mean across benchmarks present in the table. CVB: CVBench; HR-4K/8K: HR-Bench 4K/8K; Zoom: ZoomBench.
General
High-resolution
Model
CVB
MME-RW
BLINK
GQA
CountQA
ChartQA
DocVQA
OCRB
EvoChart
HallB
AMBER
V ∗
HR-4K
HR-8K
Zoom
Avg.
Qwen3.5-4B
89.16
70.30
71.23
69.32
46.86
86.28
96.45
87.30
77.36
69.82
88.80
82.72
88.12
81.50
50.65
77.06
Vision-OPD
88.51 − 0.65
75.26 + 4.96
67.75 − 3.48
68.23 − 1.09
42.74 − 4.12
86.04 − 0.24
96.24 − 0.21
87.00 − 0.30
75.52 − 1.84
67.09 − 2.73
86.32 − 2.48
92.15 + 9.43
83.50 − 4.62
81.00 − 0.50
65.21 + 14.56
77.50 + 0.45
Where-OPD
90.45 + 1.29
74.59 + 4.29
71.23 ± 0.00
75.39 + 6.07
54.12 + 7.26
86.56 + 0.28
97.01 + 0.56
89.20 + 1.90
78.64 + 1.28
81.60 + 11.78
94.37 + 5.57
88.48 + 5.76
86.62 − 1.50
83.25 + 1.75
53.61 + 2.96
80.34 + 3.28
Table 3: Best tested inference settings on Qwen3.5-4B. For each model and benchmark, the reported accuracy (%) is the best across thinking enabled or thinking disabled. Settings are selected separately for each cell; Avg. is the mean of the 15 selected scores. Smaller numbers show percentage-point changes from the base model. CVB: CVBench; MME-RW: MME-RealWorld; OCRB: OCRBench; HallB: HallusionBench; HR-4K/8K: HR-Bench 4K/8K; Zoom: ZoomBench.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Category
n
Qwen3.5-4B
Where-OPD
Δ
CVBench
Count
788
72.72
76.65
+3.93
Depth
600
95.50
96.33
+0.83
Distance
600
90.17
93.50
+3.33
Relation
650
95.38
96.77
+1.39
Overall
2,638
87.45
89.92
+2.47
Appendix
Table 4: Per-category accuracy on Qwen3.5-4B. Accuracy (%) for Qwen3.5-4B and Where-OPD across six benchmarks, broken down by their category labels. n is the number of examples in each category, and Δ is Where-OPD minus base accuracy. Bold indicates the higher score.
Benchmarks
CVB
V ∗
HR-4K
HR-8K
Zoom
CountQA
EvoChart
Avg.
Qwen3.5-4B base model
87.45
82.72
86.25
80.00
50.65
32.00
68.32
69.63
Training-scene resolution
10242 (ours)
89.92
88.48
85.50
83.25
51.01
33.77
78.64
72.94
7682
89.99
86.39
85.12
79.50
49.94
33.57
77.52
71.72
5122
88.78
85.86
84.12
81.12
50.41
34.69
78.64
71.95
Appendix
Table 5: Scene-generation ablations on Qwen3.5-4B. Top: the same scenes and questions rendered at different resolutions, with hint coordinates rescaled accordingly. Bottom: different ranges for the number of objects per scene. Results are accuracy (%); Avg. is the unweighted mean across seven benchmarks. The highlighted rows show the default setting in both blocks; bold marks the best result within each block. CVB: CVBench; HR-4K/8K: HR-Bench 4K/8K; Zoom: ZoomBench.
Benchmarks
Examples
Steps
CVB
V ∗
HR-4K
HR-8K
Zoom
CountQA
EvoChart
Avg.
Qwen3.5-4B
87.45
82.72
86.25
80.00
50.65
32.00
68.32
69.63
1,000
10
87.00
84.29
83.50
79.00
47.93
34.42
70.24
69.48
2,000
20
89.54
86.91
85.88
81.88
50.65
33.97
78.24
72.44
3,000
31
89.92
88.48
85.50
83.25
51.01
33.77
78.64
72.94
6,000
62
88.97
83.77
84.12
81.50
49.94
35.86
79.20
71.91
Appendix
Table 6: Training-set size. Accuracy (%) for Qwen3.5-4B (top) and Qwen3-VL-4B (bottom) after one epoch at each dataset size. The highlighted rows are the setting used in the main experiments. Avg. is the unweighted mean across seven benchmarks; bold marks the best dataset size within each backbone. CVB: CVBench; HR-4K/8K: HR-Bench 4K/8K; Zoom: ZoomBench.
Benchmarks
Teacher update rate η
CVB
V ∗
HR-4K
HR-8K
Zoom
CountQA
EvoChart
Avg.
Qwen3.5-4B
87.45
82.72
86.25
80.00
50.65
32.00
68.32
69.63
0 (frozen)
89.92
88.48
85.50
83.25
51.01
33.77
78.64
72.94
0.0001
89.61
89.01
83.12
81.50
51.36
35.47
77.92
72.57
0.001
89.31
86.91
84.75
80.38
51.01
34.75
77.84
72.14
0.01
89.27
84.82
84.25
81.62
51.01
34.75
76.88
71.80
Appendix
Table 7: Teacher update-rate study. Accuracy (%) for Qwen3.5-4B (top) and Qwen3-VL-4B (bottom) at different EMA update rates η ; η=0 leaves the teacher frozen. The highlighted rows are used in the main experiments. Avg. is the unweighted mean across seven benchmarks; bold marks the best update rate within each backbone. CVB: CVBench; HR-4K/8K: HR-Bench 4K/8K; Zoom: ZoomBench.
General
High-resolution
Model
CVB
MME-RW
BLINK
GQA
CountQA
ChartQA
DocVQA
OCRB
EvoChart
HallB
AMBER
V ∗
HR-4K
HR-8K
Zoom
Avg.
Qwen3.5-4B base model
89.16
70.30
71.23
64.86
46.86
86.28
96.45
86.90
77.36
69.82
88.61
82.72
88.12
81.50
46.15
76.42
Vision-OPD
88.51 − 0.65
70.43 + 0.13
67.75 − 3.48
62.26 − 2.60
42.74 − 4.12
86.04 − 0.24
96.24 − 0.21
85.80 − 1.10
75.52 − 1.84
67.09 − 2.73
83.88 − 4.73
92.15 + 9.43
83.50 − 4.62
81.00 − 0.50
57.99 + 11.84
76.06 − 0.36
OPD-V
86.35 − 2.81
69.81 − 0.49
66.12 − 5.11
58.45 − 6.41
45.03 − 1.83
84.04 − 2.24
96.09 − 0.36
84.70 − 2.20
74.40 − 2.96
64.25 − 5.57
78.13 − 10.48
89.53 + 6.81
84.25 − 3.87
79.75 − 1.75
57.28 + 11.13
74.55 − 1.88
S2VOPD
88.82 − 0.34
71.38 + 1.08
70.65 − 0.58
61.71 − 3.15
53.93 + 7.07
85.04 − 1.24
96.28 − 0.17
86.00 − 0.90
75.20 − 2.16
66.56 − 3.26
83.91 − 4.70
85.34 + 2.62
86.75 − 1.37
83.50 + 2.00
53.61 + 7.46
76.58 + 0.16
Where-OPD
90.45 + 1.29
71.50 + 1.20
71.23 ± 0.00
65.81 + 0.95
54.12 + 7.26
86.16 − 0.12
96.49 + 0.04
86.60 − 0.30
76.00 − 1.36
69.40 − 0.42
87.61 − 1.00
83.25 + 0.53
86.62 − 1.50
81.75 + 0.25
53.61 + 7.46
77.37 + 0.95
Appendix
Table 8: Evaluations under thinking enabled mode on Qwen3.5-4B. All models are evaluated with thinking enabled. Results are accuracy (%); Avg. is the unweighted mean across 15 benchmarks. The smaller numbers below each score show the percentage-point change from the base model evaluated in the same mode. CVB: CVBench; MME-RW: MME-RealWorld; OCRB: OCRBench; HallB: HallusionBench; HR-4K/8K: HR-Bench 4K/8K; Zoom: ZoomBench.
Model
L3
L7
L11
L15
L19
L23
L27
L31
Chosen
Qwen3.5-4B
0.68
0.27
0.64
0.94
0.92
0.76
1.15
0.54
L27
Vision-OPD
1.12
0.71
3.01
4.75
3.20
2.31
2.57
1.47
L15
Where-OPD
1.32
1.17
3.36
5.79
4.30
2.69
2.29
1.42
L15
Appendix
Table 9: Attention-layer selection on synthetic counting scenes. Mean enrichment of attention on the target objects at each full-attention layer, measured on a separate selection split of 158 held-out images with at least three targets. Enrichment of 1 means the targets receive attention proportional to their area. Chosen is the layer with the highest enrichment for each model. The base model is Qwen3.5-4B.
Multimodal Large Language Models (MLLMs) still struggle with fine-grained visual understanding, where answers often depend on small but decisive evidence in the full image. We observe a regional-to-global perception gap: the same MLLM answers fine-grained questions more accurately when conditioned on evidence-centered crops than on the corresponding full images, suggesting that many failures stem from difficulty to focus on relevant evidence rather than insufficient local recognition ability. Motivated by this observation, we propose Vision-OPD (Vision On-Policy Distillation), a regional-to-global self-distillation framework that transfers the model's own privileged regional perception to its full-image policy. Vision-OPD instantiates two conditional policies from the same MLLM: a crop-conditioned teacher and a full-image-conditioned student. The student generates on-policy rollouts, and Vision-OPD minimizes token-level divergence between the teacher and student next-token distributions along these rollouts. This enables the model to internalize the benefit of visual zooming without external teacher models, ground-truth labels, reward verifiers, or inference-time tool use. Experiments on multiple fine-grained visual understanding benchmarks show that Vision-OPD models achieve competitive or superior performance against much larger open-source, closed-source, and "Thinking-with-Images" agentic models. The code is available at https://github.com/VisionOPD/Vision-OPD
Qianhao Yuan, Jie Lou, Xing Yu +4
Chinese Information Processing Laboratory, Institute of Software, Chinese Academy of Sciences · University of Chinese Academy of Sciences · Xiaohongshu Inc.
On-policy self-distillation (OPSD) trains a model on its own rollouts and uses a frozen copy to provide dense token-level targets conditioned on a reference target. This works well for LLM reasoning, but a direct extension to multimodal large language models (MLLMs) can create a shortcut: the privileged target may guide tokens mainly based on the text reference target rather than the image. We propose ViGOS, a visually grounded OPSD framework for MLLM post-training. The student first writes a visual description and then reasons toward the final answer. For valid rollouts, an image-only perception teacher supervises the description, while a privileged reasoning teacher supervises the reasoning and final answer on the same student prefix. A reference teacher is used only for invalid rollouts to recover the output format. Across general vision-language, expert reasoning, visual math, spatial grounding, and visual-language-prior benchmarks, ViGOS keeps the main benefits of OPSD and improves image-grounded behavior in shortcut-prone settings.
Sihan Wang, Xiyao Liu, Lianqing Liu +1
State Key Laboratory of Robotics and Intelligent Systems, Shenyang Institute of Automation, Chinese Academy of Sciences
Self-improvement for multimodal large language models (MLLMs) is typically driven by reward-based methods that provide only coarse scalar feedback. Distillation offers a richer alternative through dense token-level supervision, but in the visual domain it usually depends on privileged context constructed using external annotations and tools, or stronger models. We introduce \textbf{CVPD} (Contrastive Counterfactual Visual Process Distillation), which, to the best of our knowledge, is the first fully self-contained framework for dense, on-policy, token-level visual self-distillation for MLLMs. CVPD identifies visual blind spots where zooming into a region changes and sharpens the model's answer distribution, while removing the same region leaves the full-image behavior largely unchanged. Such regions reveal perceptual information that the model can encode but fails to consistently utilize under full-image conditioning. We propose a three-gate Counterfactual Criterion that identifies these regions directly from the model's own responses and converts them into dense contrastive supervision for self-distillation. On Qwen3-VL-8B-Instruct, CVPD outperforms six self-evolving baselines across twelve benchmarks, including methods that rely on external GPT-4o supervision, without a single regression. It achieves gains of +3.60 on OCRBench, +3.38 on MMStar Fine-Grained Perception, and +3.08 on MMStar Logical Reasoning, while maintaining or improving performance on broader multimodal benchmarks.