cs.CVSep 29, 2026

Spatial-OPSD: Self-Improving Spatial Reasoning via Label-Free Self-Distillation

Authors: Zhenyu Liu, Zhangquan Chen, Keyi Chen, Mingze Sun, Xiang An, Haodong Jing, Ruqi Huang

Organizations: Tsinghua University · LMMs-Lab · Xi’an Jiaotong University

Abstract

Vision-language models (VLMs) increasingly operate in embodied and spatially grounded settings, where accurate understanding of depth, viewpoint, and three-dimensional relations is essential. However, improving spatial reasoning typically relies on ground-truth answers, answer-derived rewards, or other forms of task-specific supervision. We introduce Spatial-OPSD, a label-free self-improvement framework that instead exploits spatial structure naturally available from perception and reconstruction tools. During training, a privileged teacher receives automatically obtainable spatial priors, such as depth, reconstructed 3D relations, and camera geometry, while the student observes only the original visual-language input. On trajectories sampled by the student itself, the teacher provides dense token-level supervision, allowing the student to internalize spatial knowledge without ground-truth answer labels or privileged information at inference time. To extend this supervision beyond a single round, we adopt a round-wise recursive training scheme: the teacher remains frozen within each round to provide a stable learning target, and the improved student initializes both teacher and student in the next round, where privileged spatial priors re-establish an informative teacher--student asymmetry. This enables repeated self-improvement while avoiding a rapidly moving teacher during optimization. Across four VLM families, a single round of Spatial-OPSD consistently improves the five-benchmark average, while three rounds further push a strong spatially specialized model to the open-source frontier, achieving the highest average among the open models and the best results on three of five spatial reasoning benchmarks. Our code is available at https://github.com/vermouth599/Spatial-OPSD.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. The Art of Interrogation: Consistency Amplifies Factuality in Spatial Reasoning

    Jun 10, 2026Theo Uscidda, Marta Tintore Gazulla, Maks Ovsjanikov +2Spatial ReasoningLLM Reasoning Strategies

  2. ProSR: Process-Shaped Spatial Reasoning for Reliable Chain-of-Thought in VLMs

    May 25, 2026Jiangyang Li, Cong Wan, Changjie Wu +8Spatial ReasoningReasoning Trajectory

  3. Reinforcing Dual-Path Reasoning in Spatial Vision Language Models

    Jun 16, 2026Yatai Ji, An-Chieh Cheng, Yang Fu +13Spatial ReasoningIndoor Localization