Scene Understanding

Momentum

26 papers in the last four weeks, up 73% on the four weeks before. 0.3% of all new papers.

Jul 6Week of Sep 21

Latest papers 220

All topics
CardsList
  1. Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes

    Oct 1, 2026Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc +2Multimodal Large Language ModelsRecent Vision-Language Models

  2. SPHERE: Adaptive VR Indoor Scene Generation via LLM-Enhanced Spatial Preference Learning and Human-in-the-Loop RL

    Oct 1, 2026Hyeonmin Lee, Zheng Wei, Kyungmin Kwon +3Indoor Scene Generation3D Spatial Reasoning

  3. LiteReality-Agent: An Agentic System for Interactable 3D Indoor Scene Reconstruction

    Oct 1, 2026Zhening Huang, Yueyan Li, Johnathan Chiu +53D ReconstructionRgb-Depth Cameras

  4. AVSD-Scenes: A Dataset for Audio-Visual Description of Urban Scenes

    Oct 1, 2026Dhanunjaya Varma Devalraju, Arshdeep Singh, Mark D. PlumbleyAudio-Visual ReasoningMultimodal Dataset

  5. Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs

    Sep 30, 2026Taegeun Yang, Youngju Na, Yoonki Cho +1Photometric SupervisionCompositional Generalization

  6. ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing

    Sep 30, 2026Xinghao Chen, Xiangbo Gao, Jiongze Yu +2Video EditingText-To-Video Generation Model

  7. Reconstructing the Dynamic World: A Representation-Centric View of 4D Scene Reconstruction

    Sep 30, 2026Ziren Gong, Guo Chen, Yongjia Li +13Neural Radiance FieldsScene Understanding

  8. FAST: Flow Any Scene Transformer

    Sep 30, 2026Yongjian Zhang, Longguang Wang, Zhuo Song +3Harder Better Faster Denser Feature MatchingRecent Vision Foundation Models

  9. FOMO: Forget the Concept, Don't Miss Out on the Scene in Selective Video Unlearning

    Sep 30, 2026Łukasz Rudnik, Agnieszka Polowczyk, Alicja Polowczyk +1Generative Video ModelsExact Unlearning

  10. Composition, Not Conversation: VLMs Lose the Scene, Not the Thread

    Sep 29, 2026L. D. M. S. Sai Teja, Ufaq Khan, N. Siva Gopala Krishna +5TextvqaRecent Vision-Language Models

  11. Honeycomb: Constant-Size Scene Memory Representation for Video World Models

    Sep 29, 2026Jack Wei Lun Shi, Kaichen Zhou, Haoyu Chen +6Video World ModelsLong Videos

  12. Collision-Aware and Observation-Aligned Object-Centric Scene Reconstruction from Point Cloud

    Sep 29, 2026Yuxuan Xie, Xuan Yu, Rong Xiong +1Scene ReconstructionPoint Clouds

  13. Scene Retargeting: Learning Object Placement with Analogical Transfer

    Sep 29, 2026Minkwan Kim, Junho Kim, Seungmin Lee +2Cross-Embodiment RetargetingSpatial Grounding

  14. Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes

    Sep 29, 2026Xiaokang Ye, Siddhant Hitesh Mantri, Zimeng Chen +63D Scene Generation3D Editing

  15. LEGO-Anything: Coding Agents for 3D Scene Reconstruction

    Sep 28, 2026Xirui Li, Peng Shi, Mingwen Dong +73D ReconstructionScene Understanding

  16. VideoPhysEdit: Physical Counterfactual Video Editing via Rigid-Body Physical Scene Reconstruction

    Sep 28, 2026Conghan Yue, Yuanjie Chen, Yue Han +4Video EditingScene Understanding

  17. When VLMs Trust Context: Evaluating Scene Text Recognition under Misleading Context

    Sep 28, 2026Yuxing Cheng, Yuan Wu, Yi ChangScene Text RecognitionVision-Language Foundation Models

  18. When Does an Image Determine the Answer? Benchmarking Visual Answerability across Charts and Scenes

    Sep 28, 2026Sungguk Cha, Mintae Kim, Youngsub Han +2AnswerChart

  19. TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations

    Sep 24, 2026Ayush Jain, Sreeharsha Paruchuri, Ishita Gupta +53D Tracking3D Scene Understanding

  20. Beyond End-Task Success: How to Audit Visual Experience Retrieval in Robotics

    Sep 22, 2026Eshika Pathak, Leela KrishnaRobot SystemsTask Success Rate

  21. STAR: Scene- and Task-Aware 4D Radar Preprocessing Towards End-to-End Cognitive Radar

    Sep 21, 2026Seung-Hyun Song, Dong-Hee Paek, Seung-Hyun Kong4D RadarRadar

  22. Re2A: Situated Conversational Recommendation via Rubric-based Preference Reasoning and Alignment

    Sep 16, 2026Dongding Lin, Jian Wang, Xiaoyan Zhao +1RecommendationPreference Alignment Learning

  23. Online Geometric Change Detection via Scene Decomposition

    Sep 15, 2026David Thorne, Samuel Jia Cong Chua, Nakul Joshi +4Change DetectionSimultaneous Localization And Mapping

  24. SceneBench: A Hierarchical Benchmark for Vision-Language Understanding of 3D Scenes

    Sep 14, 2026Anubhav Khanal, Prabigya Acharya, Roshni Poudel +53D Spatial Reasoning3D Generation

  25. ESG: Generating Physically Consistent Dynamic 3D Scenes from Text Descriptions

    Sep 14, 2026Xintong Fang, Zhiyuan Fang, Rengan Xie +63D Scene GenerationSpatio-Temporal Scene Graphs

  26. Partition-Invariant Tuning for 3D Scene Understanding

    Sep 14, 2026Hongqiang Lin, Tianle Wang, Shuiwang Li +4Point Clouds3D Scene Understanding

  27. CoralscapesV2: Panoptic and Fine-Grained Visual Scene Understanding in Coral Reefs

    Sep 14, 2026Jonathan Sauder, Thomas Ruckli, Gabrielė Strodomskytė +17CoralPlankton Classifier

  28. GeomVLA: Unifying Scene, Motion, and Action in 3D

    Sep 12, 2026Ziyin Xiong, Nikolaos Gkanatsios, Moritz Reuss +1Perception PipelineHuman Motion Prediction

  29. GoDeep: Annotation-Free Open-Vocabulary 3D Scene Understanding via Language-Space Lifting

    Sep 8, 2026Thodoris Betsas, Anastasios Doulamis, Andreas Georgopoulos3D Scene UnderstandingVisual Embeddings

  30. Spheriverse: 3D Scene Understanding from Spherical Observations in the Wild

    Sep 8, 2026Fei Teng, Sheng Wu, Mengfei Duan +73D Semantic Occupancy Prediction3D Scene Understanding

  31. FIRE3D: Feed-forward Interactive 3D Scene Reconstruction Within A Minute

    Sep 8, 2026Hongchi Xia, Tianhang Cheng, Wei-Chiu Ma +1Feed-Forward 3D Reconstruction3D Reconstruction

  32. EdMCGS: Event-Driven Markov Chain Gaussian Splatting for Extreme-Low-Frame-Rate Dynamic Scene Reconstruction

    Sep 8, 2026Yuzhong Wang, Wenmin Wang, Xinxing YuScene Reconstruction3D Gaussian

  33. CS-CLIP: Compositional Scene Graph-guided CLIP for Robust Compositional Reasoning

    Sep 8, 2026SeongJun Jeong, Minjoon Jung, Woo Suk Choi +2Scene Understanding

  34. Geometry-Informed Distributed Acoustic Scene Understanding

    Sep 7, 2026Yiyuan Yang, Shitong Xu, Niki Trigoni +1Microphone ArrayIndoor Environments

  35. Scene Graph-Driven Haptic Feedback for Safety Enhancement in Robotic Ophthalmic Surgery via Physically Simulated iOCT

    Sep 7, 2026Danial Arbabi, Korab Hoxha, Angelo Henriques +2Surgical RobotsHaptic Feedback

  36. HELIOS: From midnight to noon, continuous outdoor urban scene relighting

    Sep 1, 2026Hala Djeghim, Nathan Piasco, Luis Roldão +4Image RelightingUnpaired Image-To-Image Translation

  37. DSG: Dynamic 3D Scene Graph Construction for Embodied Agents in Changing Indoor Environments

    Sep 1, 2026Ming Liao, Chao Ye, Jianing Fei +13D Scene GraphsSpatio-Temporal Scene Graphs

  38. Scene Graph-based Driving Scenario Extraction for Automotive Egocentric Datasets

    Aug 31, 2026Stefan Ramdhan, Kyanna Dagenais, Vera Pantelic +2Naturalistic Driving DataAutomated Vehicles

  39. Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling

    Aug 31, 2026Minghan Qin, Yuang Wang, Xiuyu Yang +6Indoor Scene GenerationScene Understanding

  40. CapFrame: Text-Instructed Viewpoint Grounding in 3D Gaussian Scenes via Geometric Pseudo Labels

    Aug 31, 2026Jirong Li, Satoshi Ikehata, Shuhei Kurita +13D GaussianNovel View Synthesis

  41. ScenePilot: Grow-and-Repair Policy for Text-Driven 3D Indoor Scene Generation

    Aug 31, 2026Jiawei Zhang, Hongsong Wang, Pan Zhou3D Scene GenerationIndoor Scene Generation

  42. MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching

    Aug 12, 2026Xingwei Sun, Heinrich Dinkel, Gang Li +7Text-To-AudioSeed-Tts-Eval Benchmark

  43. InterPruner: Interactive Structured Pruning via Taylor-Implicit Criterion and Language-Prior Modulator for Multimodal Object Detection

    Aug 11, 2026Qi Ming, Zihan Yang, Shaoguang Huang +6Rgb-Thermal Object DetectionImplicit

  44. Rethinking Data Efficiency in Industrial Dense Prediction: Pretraining Coherence, Not Inductive Bias, Determines ViTs Low-Data Advantage

    Aug 11, 2026Haoran Sui, Yaoyuan JiaVision TransformerFrozen Backbone

  45. Diffuse the object, keep its label: curating detector training data from a few unlabeled photographs via VLM-built 3D vegetation scenes

    Aug 10, 2026Mario Malizia, Marnix Enting, Rob Haelterman +1Parameter-Efficient AdaptationScene Understanding

  46. SonicWeave: Chunk-Routed Mixture-of-Experts for Unified Audio Scene Generation

    Aug 10, 2026Yunrui Cai, Xu Li, Yucheng Zhou +8Modern Generative Audio ModelsAudio Understanding

  47. SeqLoc: Beyond the Single Frame for Cross-View Geo-Localization in Feature-Sparse Scenes

    Aug 8, 2026Junwei Zheng, Yun Huang, Ruize Dai +8Cross-View Geo-LocalizationOmnidirectional Images

  48. Scenix: Sparse-View 3D Scene Reconstruction via Executable Scene Programs

    Aug 7, 2026Kai Li, Lutao Jiang, Zhenyang Li +10Scene ReconstructionScene Understanding

  49. UQ-Loc: Uncertainty-Aware LiDAR Scene Coordinate Regression

    Aug 6, 2026Jacek KomorowskiScene Understanding

  50. iARCS: Iterative Agentic RL for Controllable 3D Scene Generation

    Aug 6, 2026Saugat Adhikari, Ashok Prasad Neupane, Pramish Paudel +23D Scene GenerationSynthetic-To-Real Domain Gap

  51. SR-JEPA: Learning Predictive Latent State in 3D Scenes

    Aug 6, 2026Zihan Zhou, Qifu Wen, Xi ZengNext-State PredictionLatent Variable

  52. DynaPix: Can Vision-Language Models Identify the Exact Future?

    Aug 6, 2026Thong Nguyen, Vinh-Hien Do, Quynh Vo +2Dense Temporal AnnotationScene Understanding

  53. Geospatial-Prior Guidance for 3D Semantic Scene Completion

    Aug 4, 2026Meng Wang, Shougao Zhang, Wenzhe He +4Spatial PriorsSatellite Imagery