cs.LGOct 8, 2026

SpatialOPSD: Self-Distilling Spatial Intelligence from Verified Coding Agent Traces

Authors: Rongxue Li, Meng Yang, Yiru Mao, Yongliang Tao, Lulu Hu, Bin Yang, Zhao Xu, Weihua Luo, +1 more

Organizations: Alibaba Group

Abstract

Spatial coding agents significantly improve spatial reasoning in Multimodal Large Language Models (MLLMs) by using external tools to generate verified execution traces. However, this paradigm inherently suffers from prohibitive inference-time overhead and external dependencies. In this paper, we explore whether an MLLM can internalize this agentic capability to operate entirely tool-free. We begin with a simple observation: prompting an MLLM with summarized execution traces of a spatial coding agent naturally unlocks the model's internal spatial Chain-of-Thought (CoT). Motivated by this, we introduce SpatialOPSD, an on-policy self-distillation framework that internalizes spatial reasoning into a standalone MLLM by formulating verified agent traces as privileged information. To mitigate privileged-information leakage during distillation, we introduce Repetition-Aware Distillation, which combines repetition masking with unlikelihood regularization. Experiments across multiple benchmarks demonstrate that self-distilling SpatialOPSD achieves higher average accuracy than SFT and GRPO on both spatial and OOD datasets, exhibiting superior performance and generalization.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes

    Oct 1, 2026Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc +2VLM DistillationMultimodal Large Language Models

  2. CoCoSI: Collaborative Cognitive Map Construction for Spatial Intelligence

    Jun 9, 2026Yiming Zhang, Ruoxuan Cao, Zhihang ZhongMultimodal Large Language ModelsMulti-Agent LLMs

  3. Spatial-OPSD: Self-Improving Spatial Reasoning via Label-Free Self-Distillation

    Sep 29, 2026Zhenyu Liu, Zhangquan Chen, Keyi Chen +4Visual Spatial Reasoning3D Spatial Reasoning