cs.CVSep 29, 2026

S4VY: Segment Anything in Feed-Forward 4D Visual Geometry

Authors: Jingdong Zhang, Xin Li, Jan Kautz, Wenping Wang, Chris Choy

Organizations: Texas A&M University, College Station · NVIDIA

Abstract

Accurate instance segmentation in dynamic scenes is important for downstream applications such as robotics and autonomous driving. Existing Segment Anything models operate primarily on 2D image or video masks and preserve identity through sequential memory, while promptable 4D instance segmentation built upon feed-forward visual geometry remains underexplored. We introduce S4VY, a Segment Anything model built on feed-forward 4D visual geometry. From a set of RGB observations, S4VY transforms shared visual-geometric features into an exhaustive set of class-agnostic 4D instance masks through a space-time query decoder, with each persistent object query binding one entity across all observations. This representation supports prompt-independent segmentation as well as point- and box- conditioned selection, without requiring a seed mask or temporal ordering. We further develop an agentic harness for natural-language grounding in the large observation space of a 4D scene. Active tree search identifies relevant frames without scanning every fixed window; a dual-stream grounder combines fine-grained VLM visual priors with geometry-consistent instance features through complementary bounding-box prediction and object-query matching; and an independent critic selects the final 4D instance mask from their predictions. Extensive experiments demonstrate state-of-the-art 4D instance segmentation and strong language-guided grounding performance under a unified evaluation spanning static and dynamic scenes.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer

    Jul 21, 2026Zhengyu Zou, Hao Li, Kuixuan Jiao +7Visual Geometry Grounded Transformer4D Reconstruction

  2. Dyna3: VLM-Guided Training-Free 4D Reconstruction via Depth Foundation Models

    Oct 1, 2026Xinhao Xiang, Weiyang Li, Zhijie Zheng +24D ReconstructionScene Reconstruction

  3. SAM-V: Geometry-Aware Segment Anything for Multi-View Instance Segmentation

    Sep 21, 2026Jiangshan Gong, Yuqun Wu, Qiqian Fu +4Amodal Segmentation3D Perception