Vision-language foundation models such as CLIP provide strong semantic representations, but their patch tokens are not directly optimized for dense metric geometry. SPACE-CLIP showed that frozen CLIP features can support monocular depth estimation through layer-group feature fusion, yet it leaves open how neighboring CLIP tokens should be combined to recover fine local structure. We present SPACE-CLIPv2, a frozen-backbone depth decoder that aggregates fixed local neighborhoods in CLIP token space. At selected decoder stages, the model samples a fixed token stencil, predicts aggregation weights, and injects the resulting response through a gated residual update. A token-space high-pass branch further preserves shallow local contrast. On NYU Depth V2, SPACE-CLIPv2 improves over a matched SPACE-CLIP baseline, while five-seed experiments consistently favor fixed over learned-offset sampling. Zero-shot iBims-1 evaluation further improves boundary and planar-geometry measures. These results support constrained local token aggregation as a practical mechanism for decoding geometry from frozen CLIP representations.
Figures & tables
Fig. 1: SPACE-CLIPv2 overview. A frozen CLIP vision encoder provides grouped patch-token maps to a trainable depth decoder. SPACE-CLIPv2 augments layer-level feature fusion with fixed local token aggregation: each selected decoder stage reads a fixed token-space neighborhood, predicts aggregation weights, and injects the aggregated response through a gated residual update. The shallow token-detail branch is included as a secondary refinement path and is held fixed in the core ablation.
Variant
Aggregation
AbsRel ↓
RMSE ↓
δ1↑
B-AbsRel ↓
Without local aggregation
–
0.123
0.428
0.853
0.122
Learned-offset aggregation
learned
0.119
0.420
0.860
0.118
Offset-free aggregation
fixed
0.119
0.417
0.863
0.118
TABLE I: Local token aggregation ablation on NYU Depth V2. B-AbsRel denotes boundary AbsRel. All rows use the same token-detail setting.
Metric
SPACE-CLIP
SPACE-CLIPv2
Change
AbsRel ↓
0.3353
0.3296
−1.70%
δ1↑
0.4164
0.4434
+6.48%
DBE-Acc. (px) ↓
5.124
4.081
−20.35%
DBE-Comp. (px) ↓
98.743
58.560
−40.7%
Flatness (cm) † ↓
18.972
17.709
−6.66%
Orientation ( ∘ ) † ↓
45.330
41.418
−8.63%
TABLE II: Zero-shot iBims-1 evaluation on all 100 official core images. Both models are trained only on NYU. Lower is better except for δ1 .
Model
AbsRel ↓
RMSE ↓
log10 ↓
δ1↑
DepthCLIP [ 31 ]
0.388
1.167
0.156
0.394
PureCLIP-Depth [ 17 ]
0.201
0.670
0.084
0.671
CaBins [ 24 ]
0.120
0.401
0.050
0.866
CLIP2Depth [ 13 ]
0.100
0.379
0.042
0.900
SPACE-CLIP [ 6 ]
0.123
0.435
0.053
0.849
SPACE-CLIPv2
0.119
0.417
0.051
0.863
TABLE III: CLIP-based monocular depth estimation on NYU Depth V2. Published methods may use different protocols; SPACE-CLIP and SPACE-CLIPv2 use our matched evaluation setting.
Fig. 2: Qualitative NYU Depth V2 examples. Each row shows RGB input, ground-truth depth, and SPACE-CLIPv2 prediction. Predictions preserve the main room layout and large planar structures, while remaining errors mainly occur near thin objects, boundaries, and invalid ground-truth regions.