cs.CVMar 24, 2026

UniFunc3D: Unified Active Spatial-Temporal Grounding for 3D Affordance Segmentation

Authors: Jiaying Lin, Dan Xu

Organizations: The Hong Kong University of Science and Technology

Abstract

Affordance segmentation in 3D scenes requires an agent to ground implicit natural-language instructions into precise masks of fine-grained interactive elements. Existing training-free methods typically rely on fragmented pipelines, which introduce visual blindness during task parsing and limit accuracy through single-scale spatial and temporal processing. We present UniFunc3D, a unified and training-free framework that treats the multimodal large language model as an active observer. By utilizing a unified MLLM backbone, UniFunc3D performs joint semantic-temporal-spatial reasoning to ground task decomposition in direct visual evidence. Our approach introduces active spatial-temporal grounding with a coarse-to-fine strategy. This allows the model to select correct video frames adaptively and focus on high-detail interactive parts while preserving the global context necessary for disambiguation. On SceneFun3D, our UniFunc3D achieves state-of-the-art performance, surpassing prior training-free methods by a large margin with a relative 59.9% mIoU improvement, and even outperforming training-based methods without any task-specific training. Code is available on our project page: https://jiaying.link/unifunc3d.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ThinkAfford: Affordance-Centric Reasoning for Fine-Grained 3D Grounding in Cluttered Scenes

    Aug 11, 2026Xinrui Lin, Sha Zhang, Shumin Wang +33D Visual GroundingAffordances

  2. SceneScaffold: Active Scene-State Construction for Unified 3D Scene Understanding

    Sep 27, 2026Xiangqi Li, Libo Huang, Jiarui Zhao +43D Scene UnderstandingState-Tracking