ElasticFit: Fit-Aware 3D Object Insertion via VLM Reasoning and Generative Adaptation
Organizations: Delft University of Technology Delft, Netherlands
Abstract
Inserting objects into existing 3D scenes requires more than selecting a plausible location: the inserted object must also fit local geometry while preserving semantic intent and physical plausibility. Although recent Vision-Language Models (VLMs) and generative models enable semantic reasoning and visual content creation, they offer limited 3D grounding and geometric control when an inserted object must fit into constrained local spaces. We introduce ElasticFit, a VLM-guided framework for fit-aware object insertion centered on a novel scene-grounded representation. Given a language instruction and rendered scene observations, ElasticFit infers structured fitting cues that specify where the object should be grounded, what volume it should occupy, how it should be oriented, and its adaptation mode (rigid placement, uniform scaling, or elastic fitting). These cues convert high-level VLM reasoning into explicit 3D constraints that condition object generation and guide downstream geometric fitting. ElasticFit then generates a scene-conditioned object prior, reconstructs it in 3D, and refines the mesh through mode-specific fitting while enforcing collision avoidance, contact consistency, and physical grounding. In fixed-asset baseline comparisons, ElasticFit improves spatial relation success from 50.8% to 69.7% and support success from 48.3% to 91.7% over the strongest baseline, while providing novel support for generative "make-it-fit" insertions in complex scenarios.
Figures & tables
| Subset | SRS | CFR | SupR | CPS |
| All tasks | 60.8% | 71.7% | 83.3% | 4.57 |
| Pose-and-contact tasks | 73.3% | 76.7% | 80.0% | 4.50 |
| Gap-fitting tasks | 48.3% | 66.7% | 86.7% | 4.63 |
| Setting | SRS | CFR | SupR | CPS | Pen. |
| ElasticFit | 47.8% | 76.7% | 83.3% | 4.63 | 1.08% |
| w/o physics settle | 46.7% | 76.7% | 80.0% | 4.60 | 1.45% |
| w/o geom.-phys. fitting | 47.8% | 36.7% | 83.3% | 4.23 | 6.09% |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Diagnostic category | Count / 25 | Explanation |
| SRS not satisfied, but visually plausible ( ) | 24 (96.0%) | These cases do not fully satisfy strict geometric predicates, but their rendered placements remain contextually plausible according to CPS. |
| parallel_to check not satisfied | 19 (76.0%) | These cases contain an unsatisfied parallel_to relation, indicating that orientation constraints are a common source of strict SRS misses in constrained regions. |
| Failure source | Count / 60 | Description |
| Directional fitting error | 2 (3.3%) | Incorrect yaw, pitch, alignment, or gap-fill direction. |
| Generation / adaptation mismatch | 6 (10.0%) | Object prior or deformation mismatches the target fitting region. |
| Concave support ambiguity | 1 (1.7%) | Sink-like concave geometry can confuse inside placement with rim or surface support. |
| Aspect | FirePlace | ElasticFit |
| Target object | Pre-existing fixed asset | Generated or fixed object |
| Main objective | Place a given object | Generative “make-it-fit” insertion |
| Adaptation | Pose-level placement | Rigid, uniform, or elastic fitting |
| Constrained spaces | Limited by fixed-asset geometry | Designed for gaps, corners, wall-side, under/inside spaces |
| Grounding | Fine-grained geometric placement constraints | Grounded fitting representation with support, ghost box, orientation, mode, scale, and prompt |
| Physical refinement | Geometry-aware collision/support refinement | Geometry-aware fitting plus physics-guided settling |
| Method | Target Object | Placement Strategy | Adaptation | Constrained-Space Insertion |
| LayoutGPT | Fixed / retrieved | Layout pose direct prediction | Rigid only | No |
| Holodeck | Fixed / retrieved | Layout constraints / box solver | Rigid only | No |
| LayoutVLM | Fixed / retrieved | Visual layout / box solver | Rigid only | No |
| FirePlace | Fixed asset | Fine-grained geom. constraints | Rigid only | Fixed-asset only |
| ElasticFit | Generated / fixed | Grounded rep. + geom.-phys. fitting | Rigid / Uniform / Elastic | Yes |
| Field | Type | Description |
| anchor_object_id | int | Main reference object for spatial reasoning. |
| support_object_id | int | Object expected to support the inserted object. |
| target_uv | Image-space target point for the object’s bottom center. | |
| look_at_uv | Optional point that the object front should face. | |
| top_uv | Optional top contact point for leaning or tilted poses. | |
| scale | Object size in the default upright pose. |
| Prompt block | Purpose |
| Role and task | Defines the VLM as a 3D spatial reasoning stage for placing a new object in an existing 3D scene. |
| Scene context and inputs | Provides image resolution, normalized image coordinates, visible object IDs, object metadata, and the user instruction. |
| Reference and scale rules | Selects a valid anchor object and predicts the target object dimensions in its default upright pose. |
| Coordinates and placement rules | Predicts the target support point, optional facing point, and optional top-contact point for leaning or lying poses. |
| Physical attributes and orientation | Predicts adaptation mode, pose tag, orientation relation, yaw offset, pitch offset, and orientation reference object. |
| Snapping and gap filling | Handles contact-heavy instructions such as against or flush with , and fill-space instructions such as fill , occupy , or fit . |
| Mode | Scale | Main variable | Solver |
| Rigid | Translation near ghost box | Heuristic search | |
| Uniform | Translation and global scale | Gradient optimization | |
| Elastic | Box-local extents | Box-aligned expansion |
| Task set | Subset / category | Primary constraint types |
| Common Placement | Distance | near, far |
| Directional | left of, right of, in front of, behind | |
| Facing | face to | |
| Alignment | aligned with, parallel to | |
| Against wall | against wall | |
| Fit-Critical | Pose-and-contact | user-defined rotation, leaning, lying |
| Method | Spatial semantics | Physical plausibility |
| LayoutGPT | Directly predicts target position and yaw from a CSS-style scene description. | No explicit collision or support refinement; placements are only clamped to scene bounds. |
| Holodeck | Generates Holodeck-style spatial constraints and searches over discrete placement candidates. | Uses AABB-level collision, support, wall, and relation checks during candidate scoring. |
| LayoutVLM | Initializes object pose from vision-language reasoning and semantic constraints. | Uses bounding-box-level differentiable optimization for feasible placement refinement. |
| ElasticFit | Uses view-aligned structured grounding with support validation and local scene cues. | Uses geometry-aware local mesh grounding and physics-guided placement refinement. |
| Prompt block | Purpose |
| Task instruction | Specifies single-object insertion into a fixed existing 3D scene. |
| Coordinate system | Defines left , depth , top , and yaw orientation. |
| Scene bounds | Provides the valid 3D range of the scene in centimeters. |
| Output template | Enforces one CSS-style block for the target object. |
| Rules | Predicts only position and yaw while keeping the target asset fixed. |
| Scene objects | Lists context objects with captions, size, position, and orientation. |
| Method | Avg. total time (s) | Notes |
| LayoutGPT | 42.97 | End-to-end runtime |
| Holodeck | 16.71 | End-to-end runtime |
| LayoutVLM | 100.03 | End-to-end runtime |
| ElasticFit | 11.53 | Placement-only; 10.50 s VLM + 1.03 s fitting |
| Assets | ||
| Alarm clock | Book | Chair |
| Backpack | Bowl | Coffee machine |
| Bike | Camera bag | Fan |
| Camera | Cap | Glasses |
| Lamp | Laptop | Monitor |
| Mug | Picture frame | Pillow |