3D Scene Understanding

Momentum

27 papers in the last four weeks, up 200% on the four weeks before. 0.3% of all new papers.

Jul 13Week of Sep 28

Latest papers 188

All topics
CardsList
  1. ECHO: Embodied Camera Observations of Human Object Carrying

    Oct 7, 2026Xuefei Sun, Lorin Achey, Kali Hamilton +4Human-Object Interaction3D Scene Understanding

  2. M3SunAgent: Monocular 3D Spatial Understanding Agent for Metric Depth Estimation and 3D Visual Grounding

    Oct 6, 2026Jinsong Zhang, Kejun Wu, Ming Zhu +3Monocular Depth EstimationTool-Using Agents

  3. GS-Pool: Object-Level Change Detection in 3D Gaussian Splatting

    Oct 5, 2026Boaz Keren-Gil, James Gain, Patrick Marais3D Gaussian Splatting3D Scene Understanding

  4. ArticuTable: Generating Instance-Level Interactive Rigid-Articulated 3D Tabletop Scenes from a Single Image

    Oct 4, 2026Kai Lv, Yibo Yin, Lijun Guo +3Articulated Object Reconstruction3D Reconstruction

  5. Task-Adaptive Grounded 3D-Programmers Using 2D VLMs

    Oct 1, 2026Arman Raayatsanati, Sombit Dey, Anna-Maria Halacheva +33D Scene Generation3D Scene Understanding

  6. WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents

    Sep 30, 2026Ziyan Jiang, Jingbo Yang, Jiabao Ji +5Multimodal Anomaly DetectionMultimodal Agents

  7. STARS: From Spatiotemporal Dynamics to Social Representations in Human-Robot Interaction

    Sep 30, 2026Nathan Tsoi, Michael J. Munje, Tejas Oberoi +5VLM EvaluationSocial Robot Navigation

  8. HIGS: Hierarchical Implicit Grids for Joint Geometric and Semantic Scene Understanding

    Sep 29, 2026Hanwen Cao, Wenqiang Wu, Kuang-Ting Tu +5Implicit Neural RepresentationsHierarchical Representation Learning

  9. Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

    Sep 29, 2026Jaewoo Jung, Hyeonseo Yu, Honggyu An +103D Spatial Reasoning3D Scene Representation

  10. UniAfford: Token-Routed Multitask Learning for Generalizable 2D-3D Affordance Perception

    Sep 29, 2026Yuhao Liu, Yiming Zhong, Hanqing Wang +9Cross-Modal Representation LearningUnified Multimodal Models

  11. ExcavaTwin: Training-Free Geometry-Guided Semantic Elevation Mapping for Autonomous Excavation

    Sep 28, 2026Yu Deng, Lingshan Zeng, Tong Hu +1Semantic Mapping3D Scene Understanding

  12. SceneScaffold: Active Scene-State Construction for Unified 3D Scene Understanding

    Sep 27, 2026Xiangqi Li, Libo Huang, Jiarui Zhao +43D Scene Representation3D VQA

  13. ReLoc: Rethinking Scene Coordinate Regression Architecture for Robust Outdoor LiDAR-based Localization

    Sep 27, 2026Heejoon Moon, Yurim Cho, Je Hyeong HongPoint Cloud LearningLiDAR Localization

  14. PlenoCI: Plenoptic CharacterIstics for View Dependence Aware Change Classification

    Sep 24, 2026Jason Lai, Chamuditha Jayanga Galappaththige, Niko Suenderhauf +23D Gaussian SplattingSemantic Change Detection

  15. VLMs Can Describe, But Not Measure: Object-Centric Scene Understanding for Robotic Manipulation

    Sep 23, 2026Enrico Saccon, Tommaso Faraci, Iñigo De La Ossa Zarzuelo +3Robotic ManipulationVLMs for Robotics

  16. CODA: Depth-Aligned Scene Completion and Object Decomposition from a Single RGB-D Image

    Sep 22, 2026Dongwon Son, Junhyek Han, Yoontae Cho +53D ReconstructionIndoor Scene Reconstruction

  17. Geometric and Semantic Coupling for Interaction Understanding in 3D Scenes

    Sep 21, 2026Hanyang Kong, Xingyi Yang3D Part Segmentation3D Scene Understanding

  18. SenseFuse: Label-Free Fusion of Image and Shape Encoders for Open-Vocabulary 3D Instance Segmentation

    Sep 17, 2026Euiseok Han, Tri Ton, Hwanhee Kim +2Open-Vocabulary 3D SegmentationMultimodal Fusion

  19. SSC-Priors: Exploring Semantic and Visibility Priors to Boost Lidar Semantic Scene Completion

    Sep 15, 2026Tetiana Martyniuk, Jonathan Seele, Alexandre Boulch +33D Scene Completion3D Scene Understanding

  20. Exploring 2D backbone effects for indoor semantic occupancy prediction

    Sep 15, 2026Shizhang Fanga, Wanling Yea, Qi ZhengVision Foundation ModelsSemantic Occupancy Mapping

  21. NeuroSymbEAD: A Large Scale Neuro-Symbolic Caption Dataset for Omni-Directional Embodied Autonomous Driving

    Sep 15, 2026Muhammad Ahmed Ullah Khan, Mohammed Elamine, Sheikh Talha Uddin +3Autonomous Driving DatasetsAutonomous Driving Benchmarks

  22. SceneBench: A Hierarchical Benchmark for Vision-Language Understanding of 3D Scenes

    Sep 14, 2026Anubhav Khanal, Prabigya Acharya, Roshni Poudel +5Spatial Reasoning Benchmarks3D Spatial Reasoning

  23. Partition-Invariant Tuning for 3D Scene Understanding

    Sep 14, 2026Hongqiang Lin, Tianle Wang, Shuiwang Li +4Point Cloud LearningParameter-Efficient Fine-Tuning

  24. ProClosure: Hierarchical Room-Object Assignment using Progressive Boundary Closure from Monocular Video

    Sep 14, 2026Vinoth Kumar Muthuraj, Soumyadeep Banik, Kushal Sharma +13D SGG3D Scene Understanding

  25. Assisted Spatial Cognition Through Vision-Language Models

    Sep 14, 2026H. Riaz, J. B. Fernandez, I. Mills +3Visual Spatial ReasoningVision-Language Navigation

  26. SAMV-DUSt3R: Instance-Centric 3D Scene Decoupling from Sparse Multi-Views

    Sep 11, 2026Langxu Zhao, Zuan Gu, Yingdan Zhang +23D ReconstructionMulti-View 3D Reconstruction

  27. GoDeep: Annotation-Free Open-Vocabulary 3D Scene Understanding via Language-Space Lifting

    Sep 8, 2026Thodoris Betsas, Anastasios Doulamis, Andreas GeorgopoulosOpen-Vocabulary 3D Segmentation3D Semantic Segmentation

  28. Spheriverse: 3D Scene Understanding from Spherical Observations in the Wild

    Sep 8, 2026Fei Teng, Sheng Wu, Mengfei Duan +7Semantic Occupancy Mapping3D Object Detection

  29. Functional-SLAM: Interaction-Aware Mapping with Online Functional Scene Graphs

    Sep 7, 2026Xinggang Hu, Chenyangguang Zhang, Zihan Zhu +3Simultaneous Localization and MappingVisual SLAM

  30. MV-STRIDE: Enabling MLLMs to Master Multi-View Spatial Reasoning via Hierarchical Capability Modeling

    Sep 7, 2026Jin Xu, Xiaojian Huang, Zhuodong Luo +6Spatial Reasoning Benchmarks3D Spatial Reasoning