S4VY: Segment Anything in Feed-Forward 4D Visual Geometry
Organizations: Texas A&M University, College Station · NVIDIA
Abstract
Accurate instance segmentation in dynamic scenes is important for downstream applications such as robotics and autonomous driving. Existing Segment Anything models operate primarily on 2D image or video masks and preserve identity through sequential memory, while promptable 4D instance segmentation built upon feed-forward visual geometry remains underexplored. We introduce S4VY, a Segment Anything model built on feed-forward 4D visual geometry. From a set of RGB observations, S4VY transforms shared visual-geometric features into an exhaustive set of class-agnostic 4D instance masks through a space-time query decoder, with each persistent object query binding one entity across all observations. This representation supports prompt-independent segmentation as well as point- and box- conditioned selection, without requiring a seed mask or temporal ordering. We further develop an agentic harness for natural-language grounding in the large observation space of a 4D scene. Active tree search identifies relevant frames without scanning every fixed window; a dual-stream grounder combines fine-grained VLM visual priors with geometry-consistent instance features through complementary bounding-box prediction and object-query matching; and an independent critic selects the final 4D instance mask from their predictions. Extensive experiments demonstrate state-of-the-art 4D instance segmentation and strong language-guided grounding performance under a unified evaluation spanning static and dynamic scenes.
Figures & tables
| 1: | . |
| 2: | aggregate and decode both prediction streams in . |
| 3: | ; return . |
| function Search ( ) | |
| 4: | Uniformly sample indices from ; use every index if . |
| 5: | ; add the predictions to . |
| 6: | if the indices in are adjacent, or , then return . |
| Dataset | Supervision | Motion | Scene | Origin | Scenes/clips | Frames |
| Infinigen ( Raistrick et al., 2023 ) | Exhaustive | Static | Indoor | Synthetic | 1,466 | 146,034 |
| ScanNet++ ( Yeshwanth et al., 2023 ) | Exhaustive | Static | Indoor | Real | 907 | 760,226 |
| RealEstate10K ( Zhou et al., 2018 ) | Exhaustive | Static | Mixed | Real | 5,127 | 593,212 |
| VIPSeg ( Miao et al., 2022 ) | Exhaustive | Dynamic | Mixed | Real | 2,806 | 66,767 |
| Cityscapes-VPS ( Kim et al., 2020 ) | Exhaustive | Dynamic | Outdoor | Real | 394 | 1,970 |
| KITTI-STEP ( Weber et al., 2021 ) | Exhaustive | Dynamic | Outdoor | Real | 12 | 5,027 |
| Method | ScanNet | DAVIS | LVOS | VIPSeg | Latency | |||||||
| T - mIoU | T - SR | 4D - IoU | 4D - SR | 4D - IoU | 4D - SR | STQ | mAP | @100 | ||||
| SAM3D ( Yang et al., 2023b ) | 0.686 | 0.296 | 0.574 | – | – | – | – | – | – | – | – | – |
| DEVA ( Cheng et al., 2023 ) | 0.180 | 0.029 | 0.219 | 0.531 | 0.584 | 0.531 | 0.373 | 0.312 | 0.412 | 0.583 | 0.047 | 32.3 |
| GLEE ( Wu et al., 2024 ) | 0.127 | 0.038 | 0.178 | 0.710 | 0.870 | 0.754 | 0.329 | 0.225 | 0.367 | 0.207 | 0.214 | 36.8 |
| OMG-Seg ( Li et al., 2024b ) | 0.189 | 0.067 | 0.218 | 0.632 | 0.789 | 0.691 | 0.167 | 0.102 | 0.191 | 0.554 | 0.205 | 13.7 |
| SAM 2 ( Ravi et al., 2025 ) | 0.635 | 0.438 | 0.588 | 0.703 | 0.823 | 0.750 | 0.450 | 0.460 | 0.444 | 0.700 | 0.249 | 175 |
| Method | HOI4D | VISOR | ||||
| T - mIoU | T - SR | T - mIoU | T - SR | |||
| DEVA ( Cheng et al., 2023 ) | 0.594 | 0.557 | 0.250 | 0.213 | 0.111 | 0.018 |
| GLEE ( Wu et al., 2024 ) | 0.427 | 0.404 | 0.261 | – | – | – |
| OMG-Seg ( Li et al., 2024b ) | 0.272 | 0.248 | 0.192 | 0.020 | 0.016 | 0.000 |
| SAM 2 ( Ravi et al., 2025 ) | 0.572 | 0.542 | 0.219 | 0.295 | 0.176 | 0.060 |
| PanSt3R ( Zust et al., 2025 ) | 0.523 | 0.524 | 0.352 | 0.195 | 0.142 | 0.044 |
| Corruption | ScanNet | DAVIS | |||||||||
| Box | Point | Box | Point | ||||||||
| S4VY | SAM 2 | S4VY | SAM 2 | SAM-V | S4VY | SAM 2 | S4VY | SAM 2 | SAM-V | ||
| 1 | 0% | 0.806 | 0.753 | 0.793 | 0.649 | 0.567 | 0.790 | 0.809 | 0.793 | 0.771 | 0.134 |
| 2 | 0% | 0.804 | 0.844 | 0.825 | 0.713 | 0.598 | 0.789 | 0.818 | 0.802 | 0.764 | 0.124 |
| 4 | 0% | 0.800 | 0.877 | 0.825 | 0.760 | 0.589 | 0.779 | 0.827 | 0.805 | 0.760 | 0.145 |
| 8 | 0% | 0.801 | 0.878 | 0.825 | 0.769 | 0.589 | 0.807 | 0.829 | 0.810 | 0.780 | 0.129 |
| Method | VLM Size | ScanRefer | ScanNet++ | Ref-DAVIS | MeViS | ||||
| 3D mIoU | Box@.25 | Mask@.25 | 3D mIoU | Box@.25 | Mask@.25 | J&F | J&F | ||
| VLM-Grounder ( Xu et al., 2024 ) | 30.5B | 0.105 | 26.0 | 17.0 | 0.063 | 12.4 | 9.9 | 0.183 | 0.141 |
| VLM SAM 2 ( Ravi et al., 2025 ) | 30.5B | 0.308 | 34.9 | 42.6 | 0.170 | 15.1 | 26.2 | 0.709 | 0.483 |
| Sa2VA-8B ( Yuan et al., 2026a ) | 8B | 0.124 | 14.9 | 17.9 | 0.040 | 3.4 | 6.4 | 0.732 | 0.542 |
| SAM 3 ( Carion et al., 2026 ) | N/A | 0.247 | 29.7 | 33.3 | 0.172 | 19.5 | 26.3 | 0.567 | 0.369 |
| MVGGT ( Wu et al., 2026a ) | N/A | 0.325 | 35.3 | 50.1 | 0.090 | 7.8 | 14.8 | 0.005 | 0.003 |
| Method | Model Size | Configuration | 3D mask mIoU | Mask@0.5 | Box@0.5 |
| Molmo 2 ( Clark et al., 2026 ) | 4B | Multi-frame point tracking | 0.432 | 45.4 | 46.1 |
| Molmo 2 ( Clark et al., 2026 ) | 4B | Multi-frame point localization | 0.407 | 43.4 | 42.1 |
| Qwen3-VL ( Bai et al., 2025 ) | 4B | Single-view box | 0.316 | 32.2 | 32.9 |
| Qwen3-VL ( Bai et al., 2025 ) | 4B | Multi-view box | 0.202 | 21.1 | 20.4 |
| GLM-4.1V ( Hong et al., 2025 ) | 9B | Per-view box | 0.303 | 32.9 | 29.6 |
| LocateAnything ( Wang et al., 2026b ) | 3B | Per-view box | 0.287 | 31.6 | 31.6 |
| Method | Model Size | Precision | Recall | F1 | IoU | Selected frames (%) |
| VideoAgent ( Wang et al., 2024b ) | 235B | 0.305 | 0.108 | 0.144 | 0.082 | 9.5% |
| VideoTree ( Wang et al., 2025b ) | 235B | 0.283 | 0.562 | 0.341 | 0.221 | 56.6% |
| Vidi 1.5 ( Team et al., 2025 ) | 9B | 0.626 | 0.144 | 0.192 | 0.121 | 7.0% |
| LensWalk ( Li et al., 2026c ) | 235B | 0.343 | 0.181 | 0.192 | 0.128 | 12.6% |
| Uniform top-8 | N/A | 0.291 | 0.144 | 0.173 | 0.099 | 13.8% |
| Uniform top-16 | N/A | 0.285 | 0.274 | 0.250 | 0.151 | 27.4% |
| (a) Grounder inputs | |||
| Input | 3D mIoU | Mask@.5 | Box@.5 |
| Single w/o | 0.206 | 22.9 | 20.7 |
| Single | 0.370 | 42.0 | 37.7 |
| Multi | 0.398 | 45.2 | 40.3 |
| Multi CoT | 0.407 | 46.0 | 41.1 |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| (a) Backbone adaptation | |||
| Strategy | mAP | AP 50 | AP 75 |
| Full fine-tuning | 0.200 | 0.286 | 0.233 |
| LoRA | 0.210 | 0.287 | 0.260 |