cs.CVOct 5, 2026

ElasticFit: Fit-Aware 3D Object Insertion via VLM Reasoning and Generative Adaptation

Authors: Tzu-Hsin Hsieh, Ricardo Marroquim

Organizations: Delft University of Technology Delft, Netherlands

Abstract

Inserting objects into existing 3D scenes requires more than selecting a plausible location: the inserted object must also fit local geometry while preserving semantic intent and physical plausibility. Although recent Vision-Language Models (VLMs) and generative models enable semantic reasoning and visual content creation, they offer limited 3D grounding and geometric control when an inserted object must fit into constrained local spaces. We introduce ElasticFit, a VLM-guided framework for fit-aware object insertion centered on a novel scene-grounded representation. Given a language instruction and rendered scene observations, ElasticFit infers structured fitting cues that specify where the object should be grounded, what volume it should occupy, how it should be oriented, and its adaptation mode (rigid placement, uniform scaling, or elastic fitting). These cues convert high-level VLM reasoning into explicit 3D constraints that condition object generation and guide downstream geometric fitting. ElasticFit then generates a scene-conditioned object prior, reconstructs it in 3D, and refines the mesh through mode-specific fitting while enforcing collision avoidance, contact consistency, and physical grounding. In fixed-asset baseline comparisons, ElasticFit improves spatial relation success from 50.8% to 69.7% and support success from 48.3% to 91.7% over the strongest baseline, while providing novel support for generative "make-it-fit" insertions in complex scenarios.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation

    Nov 18, 2025Weimin Bai, Yubo Li, Weijian Luo +43D Generative Models3D Spatial Reasoning

  2. ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation

    Jul 7, 2026Tianjiao Yu, Xinzhuo Li, Yifan Shen +43D Generative Models3D Representation

  3. Global Graph-Validated Optimization for VLM-based 3D Indoor Scene Generation

    Aug 4, 2026Jialu Huang, Yingxuan You, Fei Wang +13D Scene Generation3D Scene Understanding