Vision-language models (VLMs) increasingly operate in embodied and spatially grounded settings, where accurate understanding of depth, viewpoint, and three-dimensional relations is essential. However, improving spatial reasoning typically relies on ground-truth answers, answer-derived rewards, or other forms of task-specific supervision. We introduce Spatial-OPSD, a label-free self-improvement framework that instead exploits spatial structure naturally available from perception and reconstruction tools. During training, a privileged teacher receives automatically obtainable spatial priors, such as depth, reconstructed 3D relations, and camera geometry, while the student observes only the original visual-language input. On trajectories sampled by the student itself, the teacher provides dense token-level supervision, allowing the student to internalize spatial knowledge without ground-truth answer labels or privileged information at inference time. To extend this supervision beyond a single round, we adopt a round-wise recursive training scheme: the teacher remains frozen within each round to provide a stable learning target, and the improved student initializes both teacher and student in the next round, where privileged spatial priors re-establish an informative teacher--student asymmetry. This enables repeated self-improvement while avoiding a rapidly moving teacher during optimization. Across four VLM families, a single round of Spatial-OPSD consistently improves the five-benchmark average, while three rounds further push a strong spatially specialized model to the open-source frontier, achieving the highest average among the open models and the best results on three of five spatial reasoning benchmarks. Our code is available at https://github.com/vermouth599/Spatial-OPSD.
Figures & tables
Figure 1: Spatial-OPSD in context. Left: Spatial-OPSD uses tool-enhanced spatial priors to guide a privileged teacher during training; the student infers from the original visual question without GT answers or tools. Right: Schematic comparison of fixed-answer off-policy training, one-round training, and round-wise self-evolution. Within a round, the teacher ( T ) is frozen while the student ( S ) improves; at round boundaries, both roles are initialized from the improved student, while privileged spatial context restores an informative teacher–student asymmetry.
Figure 2: Spatial-OPSD pipeline. Native metadata or external tools produce a scene scaffold. The student receives (v,q) , while the same-checkpoint teacher additionally receives a question-relevant subgraph z . Task answers are excluded from the objective; no scaffold or tool is needed at inference.
Model
Method
SPAR-Bench
MindCube-tiny
MMSI-Bench
ViewSpatial
VSI-Bench
Average
Qwen3-VL-4B ( Bai et al., 2025 )
Base
35.148
24.57
28.0
39.01
55.43
36.432
Spatial-OPSD
44.788
31.81
29.0
40.70
55.73
40.406
Gemma-3-4B ( Gemma Team, 2025 )
Base
30.640
37.04
27.4
31.90
25.86
30.568
Spatial-OPSD
36.302
40.28
27.9
24.74
25.74
30.992
InternVL3.5-4B ( Wang et al., 2025b )
Base
29.998
35.71
28.2
35.17
54.95
36.806
Spatial-OPSD
39.433
37.04
30.1
35.21
56.92
39.741
Table 1: One-round Spatial-OPSD across 4B-scale VLM families.
Model
SPAR-Bench
MindCube-tiny
MMSI-Bench
ViewSpatial
VSI-Bench
Average
Proprietary Models
Grok-4-2025-07-09 ( xAI, 2025 )
–
63.6
37.8
43.2
47.9
–
GPT-5-2025-08-07 ( OpenAI, 2025 )
49.7
56.3
41.8
45.6
55.0
49.68
Gemini-3-Pro-Preview ( Google DeepMind, 2025 )
48.7
70.9
45.2
50.4
52.5
53.54
Open-Source General Models
BAGEL-7B-MoT ( Deng et al., 2025 )
39.1
34.7
31.0
41.3
31.4
35.50
Table 2: Comparison with proprietary, general open-source, and spatially specialized models. Among open-source models, the best result is bold and the second best is underlined.
Method
Update
Round
SPAR-Bench
Base (Qwen3-VL-4B)
–
0
35.148
Step-wise RSI
Per step
1
0.000
Fixed teacher
None
1
44.788
2
44.170
3
43.168
Round-wise RSI (ours)
Per round
1
44.788
Table 3: Ablation of teacher-refresh frequency on SPAR-Bench.
Figure 3: Teacher-refresh ablation on Qwen3-VL-4B. (a,b) Reverse-KL loss and student output entropy with a teacher frozen within the round or refreshed after every step. (c) SPAR-Bench accuracy across three matched rounds for a fixed teacher and for round-wise refresh.
Figure 4: (a) Distillation-loss ablation on five benchmarks for Qwen3-VL-4B. (b) Training-set diagnostic by spatial-prior category.
Figure 5: Effect of numerical noise in privileged spatial priors for Qwen3-VL-4B. Each prior value is multiplied by 1+ϵ , with zero-mean relative error clipped at ±2p .
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Method
Use GT Answers
SPAR-Bench
MindCube-tiny
MMSI-Bench
ViewSpatial
VSI-Bench
Qwen3-VL-4B
Base
–
35.148
24.57
28.0
39.01
55.43
SFT
✓
51.426
32.05
24.0
39.30
51.43
GRPO
✓
49.068
29.05
28.4
39.51
54.64
OURS
×
44.788
31.81
29.0
40.70
55.73
Gemma-3-4B
Base
–
30.640
37.04
27.4
31.90
25.86
SFT
✓
52.898
35.62
25.4
40.53
26.67
Appendix
Table 4: Full results of all models and methods.
Hyperparameter
Value
Optimizer
AdamW
Learning rate
2×10−6
LR schedule
cosine, lrmin=10−8
Warm-up ratio / weight decay
0 / 0
Adam (β1,β2)
(0.9,0.98)
Gradient clipping
1.0
Appendix
Table 5: Shared Spatial-OPSD training hyperparameters. “Batch size” is the number of prompts per optimizer update and also the rollout batch size.
Current Large Reasoning Models (LRMs) exhibit remarkable general capabilities but significantly underperform in spatial reasoning tasks. Existing approaches treat this gap as a knowledge deficit, relying on supervised fine-tuning (SFT) to ingest labeled spatial data from external vision sources or synthetic engines. In contrast, we argue that for many tasks, spatial reasoning capabilities are already present in pre-trained LRMs but require alignment through logical coherence under geometric 2D and 3D constraints. In this work, we propose a self-supervised reinforcement learning (RL) framework that targets the internal reasoning process without requiring ground-truth annotations. By formalizing the notion of consistency verifiers -- reward functions that check for geometric and semantic consistency under transformations -- we demonstrate that models can improve their spatial reasoning abilities. We use both image transformations, like flipping, and textual transformations, like swapping the order of objects in the question, and propose a new optimal transport-based RL strategy, OT-GRPO, which is a minimal-matching variant of group relative policy optimization tailored to pairwise verifiers. We show that this label-free consistency training approaches the accuracy of models trained with ground-truth supervision and achieves similar generalization across diverse tasks and data domains.
Theo Uscidda, Marta Tintore Gazulla, Maks Ovsjanikov +2
CREST, ENSAE, Institut Polytechnique de Paris · Google Zurich · Google DeepMind
Reliable spatial reasoning remains a core bottleneck for vision-language models (VLMs). Existing mainstream training paradigms for spatial reasoning largely rely on outcome alignment or process imitation, lacking explicit constraints on the reasoning process, and therefore struggle to ensure genuine visual dependence and stable reasoning trajectories. In this paper, we construct a high-quality CoT dataset covering diverse spatial phenomena and diagnose the model's reasoning process, revealing two typical types of process degradation during reinforcement learning optimization: Spurious Grounding, which bypasses visual evidence, and Tail Instability, where uncertainty abnormally rises in the later stage of reasoning. To address these issues, we propose ProSR, a process-shaping optimization framework for spatial reasoning. Through a Counterfactual Invariance Penalty and a Tail Drift Penalty, ProSR extends the optimization objective from single answer correctness to two process-level dimensions: visual dependence and trajectory stability. Experiments on multiple complex and out-of-distribution spatial reasoning benchmarks show that ProSR improves answer accuracy while generating reasoning trajectories that are more stable and more dependent on visual evidence.
Jiangyang Li, Cong Wan, Changjie Wu +8
1Xi’an Jiaotong University · 2Amap, Alibaba Group · 4Shenzhen University +1
Spatial VLMs have made substantial progress in geometric perception, yet complex spatial reasoning requiring multi-step inference over depth, distance, and scene relations remains challenging. Moreover, different spatial queries call for fundamentally different strategies: some are best addressed through purely linguistic, step-by-step deduction, while others require explicit 3D grounding before quantitative inference. We present Dual-Path Spatial Reasoning via Reinforcement Learning for Spatial VLMs (SR-REAL), a unified framework that equips a spatial VLM with two complementary reasoning paths: Language-Only Reasoning (LOR), which performs step-by-step linguistic deduction, and Detect-Then-Reason (DTR), which detects 3D geometric cues (e.g., centers or bounding boxes) via region tokens before explicit geometric inference. SR-REAL begins with a cold-start supervised fine-tuning stage that constructs LOR and DTR chain-of-thought supervision and exposes a region-to-3D interface, followed by RL that optimizes the policy model with accuracy and format rewards; for DTR, a discrete center-based detection reward further refines geometric alignment. Across diverse spatial benchmarks, SR-REAL significantly outperforms spatial VLM baselines: (i) a single RL-trained model supports both reasoning paths, with DTR excelling in region-aware tasks through precise 3D localization and LOR enhancing general spatial reasoning; (ii) jointly training both paths fosters mutual reinforcement; (iii) high-quality, blended cold-start data is crucial for stable RL optimization; and (iv) the model generalizes across datasets and domains without per-task tuning, demonstrating positive transfer between LOR and DTR.
Yatai Ji, An-Chieh Cheng, Yang Fu +13
The University of Hong Kong · NVIDIA · University of California, San Diego