cs.CVJan 25, 2026

SPACE-CLIP: Spatial Perception via Adaptive CLIP Embeddings for Monocular Depth Estimation

Authors: Taewan Cho, Taeryang Kim, Andrew Jaeyong Choi

Organizations: School of Computing, Gachon University, Republic of Korea

Abstract

Robotic and autonomous systems need dense spatial cues, yet adding a dedicated depth estimator can duplicate visual processing already performed by a multimodal model. CLIP-based depth methods offer an alternative, but commonly rely on text-derived conditioning or backbone adaptation. We present SPACE-CLIP, a decoder-only framework for supervised monocular depth estimation with a frozen CLIP vision backbone and no text encoder at inference. A FiLM-conditioned semantic pathway combines global image context with multilevel patch features, while a structural pathway supplies separately processed spatial features to a hierarchical fusion decoder. Indoor and outdoor evaluations demonstrate depth reconstruction with this architecture, and controlled component comparisons support the contribution of the structural pathway. Layer-selection experiments and frequency interventions further characterize the structural pathway's contribution to depth reconstruction. A shared-backbone microbenchmark further illustrates the reduction in duplicated computation. SPACE-CLIP provides a modular approach to adding dense depth prediction to compatible visual perception stacks. Code is available at https://github.com/taewan2002/SPACE-CLIP.

Figures & tables

Explore similar work

CardsList
  1. DepthAgent: Towards Better Universal Depth Estimation via Sample-wise Expert Selection

    May 22, 2026Jie Zhu, Girish Chandar Ganesan, Xiaoming LiuMonocular Depth EstimationMonocular

  2. PXDepth: Pixel-Space Modeling for Structure Preserving Monocular Depth Estimation

    Aug 17, 2026Zhiyuan Yuan, Guanying Chen, Lingteng Qiu +3Monocular Depth EstimationDepth Estimation