cs.CVOct 4, 2026

SPACE-CLIPv2: Decoding Local Geometry from Frozen CLIP for Monocular Depth Estimation

Authors: Hyun Song, Taewan Cho, Kangmin Kim, Andrew Jaeyong Choi

Organizations: School of Computing, iRASC Lab., Gachon University

Abstract

Vision-language foundation models such as CLIP provide strong semantic representations, but their patch tokens are not directly optimized for dense metric geometry. SPACE-CLIP showed that frozen CLIP features can support monocular depth estimation through layer-group feature fusion, yet it leaves open how neighboring CLIP tokens should be combined to recover fine local structure. We present SPACE-CLIPv2, a frozen-backbone depth decoder that aggregates fixed local neighborhoods in CLIP token space. At selected decoder stages, the model samples a fixed token stencil, predicts aggregation weights, and injects the resulting response through a gated residual update. A token-space high-pass branch further preserves shallow local contrast. On NYU Depth V2, SPACE-CLIPv2 improves over a matched SPACE-CLIP baseline, while five-seed experiments consistently favor fixed over learned-offset sampling. Zero-shot iBims-1 evaluation further improves boundary and planar-geometry measures. These results support constrained local token aggregation as a practical mechanism for decoding geometry from frozen CLIP representations.

Figures & tables

Explore similar work

CardsList
  1. SPACE-CLIP: Spatial Perception via Adaptive CLIP Embeddings for Monocular Depth Estimation

    Jan 25, 2026Taewan Cho, Taeryang Kim, Andrew Jaeyong ChoiMonocular Depth EstimationRobotic Perception

  2. PXDepth: Pixel-Space Modeling for Structure Preserving Monocular Depth Estimation

    Aug 17, 2026Zhiyuan Yuan, Guanying Chen, Lingteng Qiu +3Monocular Depth EstimationDepth Estimation

  3. Unlocking Dense Metric Depth Estimation in VLMs

    May 15, 2026Hanxun Yu, Xuan Qu, Yuxin Wang +2Depth Estimation3D Representation