Vision-language foundation models such as CLIP provide strong semantic representations, but their patch tokens are not directly optimized for dense metric geometry. SPACE-CLIP showed that frozen CLIP features can support monocular depth estimation through layer-group feature fusion, yet it leaves open how neighboring CLIP tokens should be combined to recover fine local structure. We present SPACE-CLIPv2, a frozen-backbone depth decoder that aggregates fixed local neighborhoods in CLIP token space. At selected decoder stages, the model samples a fixed token stencil, predicts aggregation weights, and injects the resulting response through a gated residual update. A token-space high-pass branch further preserves shallow local contrast. On NYU Depth V2, SPACE-CLIPv2 improves over a matched SPACE-CLIP baseline, while five-seed experiments consistently favor fixed over learned-offset sampling. Zero-shot iBims-1 evaluation further improves boundary and planar-geometry measures. These results support constrained local token aggregation as a practical mechanism for decoding geometry from frozen CLIP representations.
Figures & tables
Fig. 1: SPACE-CLIPv2 overview. A frozen CLIP vision encoder provides grouped patch-token maps to a trainable depth decoder. SPACE-CLIPv2 augments layer-level feature fusion with fixed local token aggregation: each selected decoder stage reads a fixed token-space neighborhood, predicts aggregation weights, and injects the aggregated response through a gated residual update. The shallow token-detail branch is included as a secondary refinement path and is held fixed in the core ablation.
Variant
Aggregation
AbsRel ↓
RMSE ↓
δ1↑
B-AbsRel ↓
Without local aggregation
–
0.123
0.428
0.853
0.122
Learned-offset aggregation
learned
0.119
0.420
0.860
0.118
Offset-free aggregation
fixed
0.119
0.417
0.863
0.118
TABLE I: Local token aggregation ablation on NYU Depth V2. B-AbsRel denotes boundary AbsRel. All rows use the same token-detail setting.
Metric
SPACE-CLIP
SPACE-CLIPv2
Change
AbsRel ↓
0.3353
0.3296
−1.70%
δ1↑
0.4164
0.4434
+6.48%
DBE-Acc. (px) ↓
5.124
4.081
−20.35%
DBE-Comp. (px) ↓
98.743
58.560
−40.7%
Flatness (cm) † ↓
18.972
17.709
−6.66%
Orientation ( ∘ ) † ↓
45.330
41.418
−8.63%
TABLE II: Zero-shot iBims-1 evaluation on all 100 official core images. Both models are trained only on NYU. Lower is better except for δ1 .
Model
AbsRel ↓
RMSE ↓
log10 ↓
δ1↑
DepthCLIP [ 31 ]
0.388
1.167
0.156
0.394
PureCLIP-Depth [ 17 ]
0.201
0.670
0.084
0.671
CaBins [ 24 ]
0.120
0.401
0.050
0.866
CLIP2Depth [ 13 ]
0.100
0.379
0.042
0.900
SPACE-CLIP [ 6 ]
0.123
0.435
0.053
0.849
SPACE-CLIPv2
0.119
0.417
0.051
0.863
TABLE III: CLIP-based monocular depth estimation on NYU Depth V2. Published methods may use different protocols; SPACE-CLIP and SPACE-CLIPv2 use our matched evaluation setting.
Fig. 2: Qualitative NYU Depth V2 examples. Each row shows RGB input, ground-truth depth, and SPACE-CLIPv2 prediction. Predictions preserve the main room layout and large planar structures, while remaining errors mainly occur near thin objects, boundaries, and invalid ground-truth regions.
Robotic and autonomous systems need dense spatial cues, yet adding a dedicated depth estimator can duplicate visual processing already performed by a multimodal model. CLIP-based depth methods offer an alternative, but commonly rely on text-derived conditioning or backbone adaptation. We present SPACE-CLIP, a decoder-only framework for supervised monocular depth estimation with a frozen CLIP vision backbone and no text encoder at inference. A FiLM-conditioned semantic pathway combines global image context with multilevel patch features, while a structural pathway supplies separately processed spatial features to a hierarchical fusion decoder. Indoor and outdoor evaluations demonstrate depth reconstruction with this architecture, and controlled component comparisons support the contribution of the structural pathway. Layer-selection experiments and frequency interventions further characterize the structural pathway's contribution to depth reconstruction. A shared-backbone microbenchmark further illustrates the reduction in duplicated computation. SPACE-CLIP provides a modular approach to adding dense depth prediction to compatible visual perception stacks. Code is available at https://github.com/taewan2002/SPACE-CLIP.
Taewan Cho, Taeryang Kim, Andrew Jaeyong Choi
School of Computing, Gachon University, Republic of Korea
Recent monocular depth estimators achieve strong zero-shot generalization, yet often struggle to preserve fine-grained structures and object boundaries. We attribute this limitation to the prevalent combination of large-patch ViT encoders and convolutional decoders, as coarse tokenization can weaken pixel-level cues that upsampling cannot fully recover. To address this issue, we propose PXDepth, a discriminative monocular depth model that separates global context modeling from pixel-level depth prediction. Specifically, a large-patch ViT captures global scene context, while a pixel-space predictor composed of Context-Modulated Pixel Transformer blocks maintains high-resolution spatial representations throughout depth estimation. This design preserves fine structures and sharp boundaries without sacrificing global depth consistency. Across diverse zero-shot benchmarks, PXDepth combines faithful local geometry with competitive global depth accuracy while remaining efficient at inference. Our code and model are available at https://yuanzhy29.github.io/PXDepth-Page/.
Zhiyuan Yuan, Guanying Chen, Lingteng Qiu +3
1Sun Yat-sen University · 3CUHKSZ · 2Shenzhen-FNii
Vision-Language Models (VLMs) excel at 2D tasks such as grounding and captioning, yet remain limited in 3D understanding. A key limitation is their text-only supervision paradigm, which under-constrains fine-grained visual perception and prevents the recovery of dense geometry. Prior methods either distill geometry from external vision models, introducing error accumulation, or enable direct prediction with inefficient per-pixel query or coarse token-level outputs. In this paper, we propose DepthVLM, a simple yet effective framework that transforms a single VLM into a native dense geometry predictor while preserving its multimodal capability. By attaching a lightweight depth head to the LLM backbone and training under a unified vision-text supervision paradigm with a two-stage schedule, DepthVLM generates full-resolution depth maps alongside language outputs in a single forward pass. We further introduce a unified indoor-outdoor metric depth benchmark in a VLM-compatible format. Experiments show that DepthVLM significantly outperforms existing VLMs with higher inference efficiency, surpasses leading pure vision models, and improves complex 3D spatial reasoning, moving toward a truly unified multimodal foundation model. The project page is available at https://depthvlm.github.io/
Hanxun Yu, Xuan Qu, Yuxin Wang +2
Zhejiang University · Tencent Hunyuan LLM · HKUST +1