RelationVGGT: Visual Geometry Transformers for 3D Spatial Relation Segmentation
Organizations: Yonsei University · NVIDIA
Abstract
Recent advances in 3D reconstruction have progressed from per-scene optimization to feed-forward inference, and semantic scene understanding has followed suit -- yet existing methods remain confined to object-centric perception, neglecting spatial relations between objects. We formulate 3D spatial relation segmentation in a feed-forward, pose-free multi-view setting: given a visually specified subject and a relational text query, the model segments the target across views without receiving its category name. To this end, we propose RelationVGGT, a novel feed-forward framework that integrates semantic features from a visual foundation model with geometry-aware representations from a 3D geometry foundation model and leverages a relation transformer for subject-conditioned, cross-view relation prediction -- requiring neither per-scene optimization nor known camera poses. We additionally provide a fully automated annotation pipeline built on ScanNet++ with VLMs and LLMs, enabling scalable training data generation for this new task.
Figures & tables
| Method | Replica | LERF | ScanNet++ | Avg. | ||||
| mIoU ( ) | mAcc ( ) | mIoU ( ) | mAcc ( ) | mIoU ( ) | mAcc ( ) | mIoU ( ) | mAcc ( ) | |
| LangSplat Qin et al. (2024) | 0.191 | 0.256 | 0.159 | 0.254 | 0.074 | 0.064 | 0.141 | 0.191 |
| OpenGaussian Wu et al. (2024) | 0.349 | 0.470 | 0.256 | 0.378 | 0.162 | 0.254 | 0.255 | 0.367 |
| RelationField Koch et al. (2025) | 0.326 | 0.537 | 0.276 | 0.410 | 0.307 | 0.491 | 0.303 | 0.480 |
| RelationField † | 0.219 | 0.336 | 0.232 | 0.346 | 0.240 | 0.395 | 0.230 | 0.359 |
| Ours | 0.446 | 0.640 | 0.254 | 0.332 | 0.445 | 0.675 | 0.382 | 0.549 |
| Method | Small (40) | Medium (39) | Large (40) | |||
| mIoU ( ) | mAcc ( ) | mIoU ( ) | mAcc ( ) | mIoU ( ) | mAcc ( ) | |
| VideoLISA Bai et al. (2024) | 0.137 | 0.171 | 0.105 | 0.143 | 0.038 | 0.064 |
| SA2VA Yuan et al. (2025) | 0.267 | 0.321 | 0.225 | 0.289 | 0.192 | 0.225 |
| UniPixel Liu et al. (2025) | 0.423 | 0.496 | 0.328 | 0.451 | 0.354 | 0.475 |
| Ours | 0.417 | 0.550 | 0.404 | 0.640 | 0.366 | 0.575 |
| Design | mIoU | mAcc |
| Ours | 0.445 | 0.675 |
| Input feature variants | ||
| P.E(pointmap) DINOv2 | 0.370 | 0.568 |
| pi3 only | 0.422 | 0.579 |
| DINOv2 only | 0.370 | 0.579 |
| Architecture variants | ||
| Subject mask | Mask IoU | mIoU | mAcc |
| GT mask | 1.000 | 0.382 | 0.549 |
| SAM2, one click | 0.871 | 0.383 | 0.551 |
| SAM2, one box | 0.923 | 0.382 | 0.549 |
| Qwen3-VL + SAM2 | 0.719 | 0.357 | 0.512 |
| GroundingDINO + SAM2 | 0.587 | 0.322 | 0.454 |
| Predicate expression | mIoU | mAcc@0.25 | |
| Meaning-preserving synonyms (set A) | 119 | -0.006 | -0.006 |
| Meaning-preserving synonyms (set B) | 119 | -0.004 | -0.001 |
| Rare synonym expressions | 113 | -0.022 | -0.042 |
| Unseen synonym expressions | 119 | -0.031 | -0.040 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Stage | Relations after stage | % of Stage 1 |
| Stage 1: initial aggregation | 4,695 | 100.0 |
| Stage 2: filtering | 770 | 16.4 |
| Human judgment | Kept by Stage 2 | Removed by Stage 2 | Total |
| Correct | 506 / 421 | 1,151 / 1,173 | 1,657 / 1,594 |
| Incorrect | 257 / 198 | 2,748 / 2,494 | 3,005 / 2,692 |
| Total | 763 / 619 | 3,899 / 3,667 | 4,662 / 4,286 |
| Statistic | Mean (%) | 95% CI | Annotator 1 / 2 (%) |
| Precision after Stage 1 | 36.4 | [32.2, 39.2] | 35.5 / 37.2 |
| Precision after Stage 2 | 67.2 | [60.3, 73.5] | 66.3 / 68.0 |
| Incorrect proposals removed | 92.0 | [89.1, 94.0] | 91.4 / 92.6 |
| Correct proposals retained | 28.5 | [22.8, 37.2] | 30.5 / 26.4 |
| Method | Query inference | Per-scene optimization |
| RelationField | 288 s | 3 h |
| Ours | 26 ms | None |
| Method | Subject specification | Target category in input | Test-scene optimization | Avg. mIoU / mAcc@0.25 |
| MVGGT | Text | Yes | No | 0.153 / 0.217 |
| ReferSplat | Text | No | Yes | 0.213 / 0.263 |
| Ours | Visual mask | No | No | 0.382 / 0.549 |
| Ours | RelationField | ||||
| Target-view subset | Views | mIoU | mAcc | mIoU | mAcc |
| Subject visible | 257 | 0.440 | 0.677 | 0.318 | 0.510 |
| Subject not visible | 23 | 0.507 | 0.652 | 0.159 | 0.217 |
| Predicate substitution | mIoU | |
| Random relation from an incompatible family | 119 | |
| Relation that changes the correct target | 118 |