Vision-language models (VLMs) increasingly operate in embodied and spatially grounded settings, where accurate understanding of depth, viewpoint, and three-dimensional relations is essential. However, improving spatial reasoning typically relies on ground-truth answers, answer-derived rewards, or other forms of task-specific supervision. We introduce Spatial-OPSD, a label-free self-improvement framework that instead exploits spatial structure naturally available from perception and reconstruction tools. During training, a privileged teacher receives automatically obtainable spatial priors, such as depth, reconstructed 3D relations, and camera geometry, while the student observes only the original visual-language input. On trajectories sampled by the student itself, the teacher provides dense token-level supervision, allowing the student to internalize spatial knowledge without ground-truth answer labels or privileged information at inference time. To extend this supervision beyond a single round, we adopt a round-wise recursive training scheme: the teacher remains frozen within each round to provide a stable learning target, and the improved student initializes both teacher and student in the next round, where privileged spatial priors re-establish an informative teacher--student asymmetry. This enables repeated self-improvement while avoiding a rapidly moving teacher during optimization. Across four VLM families, a single round of Spatial-OPSD consistently improves the five-benchmark average, while three rounds further push a strong spatially specialized model to the open-source frontier, achieving the highest average among the open models and the best results on three of five spatial reasoning benchmarks. Our code is available at https://github.com/vermouth599/Spatial-OPSD.
Figures & tables
Figure 1: Spatial-OPSD in context. Left: Spatial-OPSD uses tool-enhanced spatial priors to guide a privileged teacher during training; the student infers from the original visual question without GT answers or tools. Right: Schematic comparison of fixed-answer off-policy training, one-round training, and round-wise self-evolution. Within a round, the teacher ( T ) is frozen while the student ( S ) improves; at round boundaries, both roles are initialized from the improved student, while privileged spatial context restores an informative teacher–student asymmetry.
Figure 2: Spatial-OPSD pipeline. Native metadata or external tools produce a scene scaffold. The student receives (v,q) , while the same-checkpoint teacher additionally receives a question-relevant subgraph z . Task answers are excluded from the objective; no scaffold or tool is needed at inference.
Model
Method
SPAR-Bench
MindCube-tiny
MMSI-Bench
ViewSpatial
VSI-Bench
Average
Qwen3-VL-4B ( Bai et al., 2025 )
Base
35.148
24.57
28.0
39.01
55.43
36.432
Spatial-OPSD
44.788
31.81
29.0
40.70
55.73
40.406
Gemma-3-4B ( Gemma Team, 2025 )
Base
30.640
37.04
27.4
31.90
25.86
30.568
Spatial-OPSD
36.302
40.28
27.9
24.74
25.74
30.992
InternVL3.5-4B ( Wang et al., 2025b )
Base
29.998
35.71
28.2
35.17
54.95
36.806
Spatial-OPSD
39.433
37.04
30.1
35.21
56.92
39.741
Table 1: One-round Spatial-OPSD across 4B-scale VLM families.
Model
SPAR-Bench
MindCube-tiny
MMSI-Bench
ViewSpatial
VSI-Bench
Average
Proprietary Models
Grok-4-2025-07-09 ( xAI, 2025 )
–
63.6
37.8
43.2
47.9
–
GPT-5-2025-08-07 ( OpenAI, 2025 )
49.7
56.3
41.8
45.6
55.0
49.68
Gemini-3-Pro-Preview ( Google DeepMind, 2025 )
48.7
70.9
45.2
50.4
52.5
53.54
Open-Source General Models
BAGEL-7B-MoT ( Deng et al., 2025 )
39.1
34.7
31.0
41.3
31.4
35.50
Table 2: Comparison with proprietary, general open-source, and spatially specialized models. Among open-source models, the best result is bold and the second best is underlined.
Method
Update
Round
SPAR-Bench
Base (Qwen3-VL-4B)
–
0
35.148
Step-wise RSI
Per step
1
0.000
Fixed teacher
None
1
44.788
2
44.170
3
43.168
Round-wise RSI (ours)
Per round
1
44.788
Table 3: Ablation of teacher-refresh frequency on SPAR-Bench.
Figure 3: Teacher-refresh ablation on Qwen3-VL-4B. (a,b) Reverse-KL loss and student output entropy with a teacher frozen within the round or refreshed after every step. (c) SPAR-Bench accuracy across three matched rounds for a fixed teacher and for round-wise refresh.
Figure 4: (a) Distillation-loss ablation on five benchmarks for Qwen3-VL-4B. (b) Training-set diagnostic by spatial-prior category.
Figure 5: Effect of numerical noise in privileged spatial priors for Qwen3-VL-4B. Each prior value is multiplied by 1+ϵ , with zero-mean relative error clipped at ±2p .
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Method
Use GT Answers
SPAR-Bench
MindCube-tiny
MMSI-Bench
ViewSpatial
VSI-Bench
Qwen3-VL-4B
Base
–
35.148
24.57
28.0
39.01
55.43
SFT
✓
51.426
32.05
24.0
39.30
51.43
GRPO
✓
49.068
29.05
28.4
39.51
54.64
OURS
×
44.788
31.81
29.0
40.70
55.73
Gemma-3-4B
Base
–
30.640
37.04
27.4
31.90
25.86
SFT
✓
52.898
35.62
25.4
40.53
26.67
Appendix
Table 4: Full results of all models and methods.
Hyperparameter
Value
Optimizer
AdamW
Learning rate
2×10−6
LR schedule
cosine, lrmin=10−8
Warm-up ratio / weight decay
0 / 0
Adam (β1,β2)
(0.9,0.98)
Gradient clipping
1.0
Appendix
Table 5: Shared Spatial-OPSD training hyperparameters. “Batch size” is the number of prompts per optimizer update and also the rollout batch size.