Recent advances in 3D reconstruction have progressed from per-scene optimization to feed-forward inference, and semantic scene understanding has followed suit -- yet existing methods remain confined to object-centric perception, neglecting spatial relations between objects. We formulate 3D spatial relation segmentation in a feed-forward, pose-free multi-view setting: given a visually specified subject and a relational text query, the model segments the target across views without receiving its category name. To this end, we propose RelationVGGT, a novel feed-forward framework that integrates semantic features from a visual foundation model with geometry-aware representations from a 3D geometry foundation model and leverages a relation transformer for subject-conditioned, cross-view relation prediction -- requiring neither per-scene optimization nor known camera poses. We additionally provide a fully automated annotation pipeline built on ScanNet++ with VLMs and LLMs, enabling scalable training data generation for this new task.
Figures & tables
Figure 1 : Overview of RelationVGGT. Multi-view images are encoded by frozen 2D and 3D visual foundation models and fused via an input mixer. The subject mask conditions the relation transformer, whose output is decoded into per-view relation feature maps. The target object is segmented via feature–text similarity with the BERT-encoded relational query.
Figure 2 : Overview of the relation annotation pipeline. (a) Instance masks are projected onto frames with SoM prompting to query a VLM for pairwise relations. (b) Proposals are aggregated into a scene graph and refined via geometry-aware postprocessing and LLM. (c) Example results showing multi-view consistent subject (green) and target (red) annotations.
Figure 3 : Qualitative results on the ScanNet++, Replica, LERF benchmarks. Please zoom for the detail. Predicted masks are overlaid in red on the target views.
Method
Replica
LERF
ScanNet++
Avg.
mIoU ( ↑ )
mAcc ( ↑ )
mIoU ( ↑ )
mAcc ( ↑ )
mIoU ( ↑ )
mAcc ( ↑ )
mIoU ( ↑ )
mAcc ( ↑ )
LangSplat Qin et al. (2024)
0.191
0.256
0.159
0.254
0.074
0.064
0.141
0.191
OpenGaussian Wu et al. (2024)
0.349
0.470
0.256
0.378
0.162
0.254
0.255
0.367
RelationField Koch et al. (2025)
0.326
0.537
0.276
0.410
0.307
0.491
0.303
0.480
RelationField †
0.219
0.336
0.232
0.346
0.240
0.395
0.230
0.359
Ours
0.446
0.640
0.254
0.332
0.445
0.675
0.382
0.549
Table 1: Quantitative comparison on the Replica, LERF, and ScanNet++ benchmarks. RelationField † denotes the relation-only variant without target-object retrieval; Avg. is the macro average across the three benchmarks.
Method
Small (40)
Medium (39)
Large (40)
mIoU ( ↑ )
mAcc ( ↑ )
mIoU ( ↑ )
mAcc ( ↑ )
mIoU ( ↑ )
mAcc ( ↑ )
VideoLISA Bai et al. (2024)
0.137
0.171
0.105
0.143
0.038
0.064
SA2VA Yuan et al. (2025)
0.267
0.321
0.225
0.289
0.192
0.225
UniPixel Liu et al. (2025)
0.423
0.496
0.328
0.451
0.354
0.475
Ours
0.417
0.550
0.404
0.640
0.366
0.575
Table 2: Comparison with video foundation models across camera-baseline ranges. Small, Medium, and Large denote near-equal groups of the 119 queries ordered by camera baseline. All methods are given the subject region and relation query without the target category; video baselines process the multi-view images as a video sequence.
Figure 4 : Qualitative comparison between our method and UniPixel. Our method produces view-consistent masks, whereas UniPixel fails to maintain consistency across views. Please zoom for the detail. Predicted masks are overlaid in red on the target views.
Design
mIoU ↑
mAcc ↑
Ours
0.445
0.675
Input feature variants
P.E(pointmap) + DINOv2
0.370
0.568
pi3 only
0.422
0.579
DINOv2 only
0.370
0.579
Architecture variants
Table 3: Ablation study on model design and the annotation pipeline, evaluated on ScanNet++ benchmark (40 queries) with similarity threshold τ=0.3 .
Subject mask
Mask IoU ↑
mIoU ↑
mAcc ↑
GT mask
1.000
0.382
0.549
SAM2, one click
0.871
0.383
0.551
SAM2, one box
0.923
0.382
0.549
Qwen3-VL + SAM2
0.719
0.357
0.512
GroundingDINO + SAM2
0.587
0.322
0.454
Table 4: Robustness to subject-mask specification across three benchmarks (119 queries). All variants use the same RelationVGGT checkpoint without retraining.
Predicate expression
n
Δ mIoU ↑
Δ mAcc@0.25 ↑
Meaning-preserving synonyms (set A)
119
-0.006
-0.006
Meaning-preserving synonyms (set B)
119
-0.004
-0.001
Rare synonym expressions
113
-0.022
-0.042
Unseen synonym expressions
119
-0.031
-0.040
Table 5: Robustness to predicate expressions. We report paired per-query changes relative to the original predicate queries. The two synonym sets use independently constructed meaning-preserving alternatives. Rare expressions use the least frequent suitable synonym observed in training (stem frequency ≤20 for 94/113 queries; six queries without a suitable alternative are excluded), whereas unseen expressions use synonyms whose stems never occur during training.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Distribution of the final relation predicates used for training. The horizontal axis shows the number of occurrences on a logarithmic scale; 54 additional low-frequency predicates are grouped for visualization.
Stage
Relations after stage
% of Stage 1
Stage 1: initial aggregation
4,695
100.0
Stage 2: filtering
770
16.4
Appendix
Table 6: Number of relation proposals that survive the annotation filtering pipeline. (25 scenes as samples) Stage 2 combines multi-view consistency with the geometry-based filtering described in Sec. 3.5 .
Human judgment
Kept by Stage 2
Removed by Stage 2
Total
Correct
506 / 421
1,151 / 1,173
1,657 / 1,594
Incorrect
257 / 198
2,748 / 2,494
3,005 / 2,692
Total
763 / 619
3,899 / 3,667
4,662 / 4,286
Appendix
Table 7: Human judgment versus the Stage-2 filtering decision. Counts are reported as annotator 1 / annotator 2; ambiguous judgments are excluded.
Statistic
Mean (%)
95% CI
Annotator 1 / 2 (%)
Precision after Stage 1
36.4
[32.2, 39.2]
35.5 / 37.2
Precision after Stage 2
67.2
[60.3, 73.5]
66.3 / 68.0
Incorrect proposals removed
92.0
[89.1, 94.0]
91.4 / 92.6
Correct proposals retained
28.5
[22.8, 37.2]
30.5 / 26.4
Appendix
Table 8: Human evaluation of the annotation pipeline on 25 ScanNet++ scenes. Confidence intervals are obtained with scene-level cluster bootstrap using 20,000 resamples.
Method
Query inference
Per-scene optimization
RelationField
288 s
> 3 h
Ours
26 ms
None
Appendix
Table 9: Inference-time comparison with RelationField on an NVIDIA L40S. RelationField additionally requires scene-specific optimization before query-time inference.
Figure 6: Prompt for grounded open-vocabulary relationship extraction.
Figure 7: Prompt for relation refinement.
Method
Subject specification
Target category in input
Test-scene optimization
Avg. mIoU / mAcc@0.25
MVGGT
Text
Yes
No
0.153 / 0.217
ReferSplat
Text
No
Yes
0.213 / 0.263
Ours
Visual mask
No
No
0.382 / 0.549
Appendix
Table 10: Comparison with closely related 3D referring-segmentation methods. The methods differ in their native input interfaces and optimization protocols; in particular, MVGGT receives the target category in the evaluated text template, whereas ReferSplat and our method do not.
Ours
RelationField
Target-view subset
Views
mIoU
mAcc
mIoU
mAcc
Subject visible
257
0.440
0.677
0.318
0.510
Subject not visible
23
0.507
0.652
0.159
0.217
Appendix
Table 11: Performance on ScanNet++ target views grouped by subject visibility. The subject is always specified in the reference view; visibility here refers only to the evaluated target view.
Predicate substitution
n
Δ mIoU
Random relation from an incompatible family
119
−0.158[−0.207,−0.111]
Relation that changes the correct target
118
−0.276
Appendix
Table 12: Meaning-changing predicate controls. Δ mIoU is measured relative to the original benchmark predicate while keeping the original target mask fixed. These controls therefore measure query sensitivity rather than segmentation accuracy for the substituted relation itself.
Recent advances in 3D datasets and multimodal models have greatly improved natural language 3D scene understanding. However, most 3D referring segmentation methods do not explicitly represent the observer viewpoint, making spatial relations such as "left," "right," "front," and "behind" ambiguous and difficult to evaluate. We introduce a viewpoint-aware 3D referring segmentation dataset containing 220k benchmark samples, and scalable to tens of millions of viewpoint-conditioned samples through dense viewpoint sampling. In this dataset, target objects can only be identified through observer-centric spatial relations, making viewpoint-conditioned grounding necessary. We construct the benchmark by leveraging camera poses to automatically annotate observer-centric relations (left/right, front/behind) together with viewpoint-independent relations (above/under). Using this benchmark, we evaluate several existing 3D large multimodal models in a zero-shot setting and find that current models struggle with viewpoint-dependent spatial instructions. We further study how explicit viewpoint information can be incorporated into 3D large multimodal models. We introduce a viewpoint representation that encodes camera poses and conditions the model on the observation viewpoint, improving segmentation accuracy on viewpoint-dependent relations and increasing mIoU from 0.30 to 0.47 compared to a model without viewpoint conditioning. The dataset, code, and trained models will be made publicly available upon acceptance.
Ayaka Nanri, Klara Reichard, Mert Kiray +3
Institute of Science Tokyo · Technical University of Munich · BMW Group +12
Visual Geometry Grounded Transformer (VGGT) recovers dense 3D scene structure from multi-view images in one forward pass, but quadratic cross-frame attention limits its scalability. Existing training-free accelerators reduce computation uniformly along one axis, missing layer heterogeneity. Our spectral, probing, and causal analyses reveal three regimes: shallow layers lack cross-view structure, middle layers drive cross-view alignment, and deep layers are redundant for dense geometry yet their cross-frame attention remains essential for pose. RegimeVGGT applies layer-wise U-shaped compression along two axes: Saliency-Guided Banded Merging protects geometry- and edge-salient tokens, while Selectively Protected K/V Downsampling preserves cross-frame spatial coverage and the pose-critical path through a phase-shifted spatial grid, a reference-frame anchor, and uncompressed camera/register tokens. Training-free, RegimeVGGT achieves a 6.7x speedup over VGGT* at matched reconstruction quality.
Jinhao You, Shuo Lyu, Zhuohang Lyu +5
University of Pennsylvania · University of California, Irvine · Nanyang Technological University +1
Feed-forward 3D models are commonly trained using either expensive geometric supervision or self-supervised photometric objectives, both of which provide incomplete learning signals. We introduce Vision-Language Reprojection Consistency (VLRC), a scalable auxiliary objective that exploits frozen vision-language representations as semantic multi-view supervision. Given a predicted 3D reconstruction, VLRC reprojects dense vision-language features across views and enforces feature consistency between corresponding image locations, requiring no additional 3D annotations. The objective integrates seamlessly with both self-supervised monocular reconstruction and supervised-pretrained feed-forward 3D models during unlabeled adaptation. By aligning geometry with language-grounded features, VLRC not only improves depth and camera estimation but also enables more coherent multi-view semantic fusion for open-vocabulary 3D scene understanding. Experiments on indoor and outdoor benchmarks demonstrate consistent gains in 3D reconstruction accuracy and zero-shot open-vocabulary 3D semantic segmentation.
Marwane Hariat, David Filliat, Antoine Manzanera
U2IS, ENSTA – Institut Polytechnique de Paris · Agence Minist´erielle pour l’IA de D´efense