Scene Retargeting: Learning Object Placement with Analogical Transfer
Authors: Minkwan Kim, Junho Kim, Seungmin Lee, Changwoon Choi, Young Min Kim
Organizations: Dept. of Electrical and Computer Engineering, Seoul National University · Interdisciplinary Program in Artificial Intelligence and INMC, Seoul National University
Interactive simulations of embodied AI or spatial computing applications build on realistic 3D scenes that support daily activities. However, sparse, irregular layout structures impose scene-specific physical constraints, making it hard to define a generalizable framework for generating similar functional context. We formalize Scene Retargeting as stably transferring the semantically coherent spatial organization across layouts, rather than relying on textual descriptions or pairwise relationships. Our cluster-wise transfer flexibly handles mismatched object instances and adapts to distinctive floor plans. We optimize to preserve the rich semantic context of individual clusters by respecting the spatial distribution of foundation features. We can then impose physical constraints to refine wall contacts, pairwise alignment, or clear passageways and openings. Our framework outperforms state-of-the-art methods on layout generation on the 3D-FRONT dataset, and demonstrates downstream applications including real-to-sim transfer, analogical trajectory transfer, and multi-reference composition.
Figures & tables
Figure 1: Scene Retargeting. From a single reference scene (a), our method reproduces semantically coherent layouts in target rooms with different floor plans and object instances (b). Functional groups and their internal relationships are preserved while maintaining physical plausibility.
Figure 2: Method Overview of Scene Retargeting. In stage 1 (Cluster Placement, blue), functional clusters decomposed from the reference scene are placed onto the target floor plan by the cluster placer. In stage 2 (Object Arrangement, green), object tokens cross-attend to spatial context features within the layout decoder to assign object-level poses. Stage 3 (Analogical Refinement, red) matches objects across scenes, optimizes pairwise relations and wall clearances, and enforces physical validity to produce the final 3D layout.
Figure 3: Analogical Refinement. Reference clusters (a) are placed on the target floor plan (b). After initial object arrangement, we apply graph matching followed by analogical refinement (c), yielding the final arrangement (d).
Method
Ref.
Physical Plausibility
Semantic Coherency
Overall
Fidelity
CF ↑
IB ↑
Pos. ↑
Rot. ↑
PSA ↑
iRecall ↑
Lego-Net ( Wei et al., 2023 )
-
84.7
69.3
40.4
41.7
23.3
43.2
DiffuScene ( Tang et al., 2024 )
Text
83.9
50.3
39.8
39.2
16.7
38.6
MiDiffusion ( Hu et al., 2026 )
-
86.7
63.7
44.7
43.6
25.7
41.4
InstructScene ( Lin and Mu, 2024 )
Text
88.4
52.7
56.5
55.1
24.8
46.9
Holodeck ( Yang et al., 2024b )
Text
77.3
64.0
43.9
44.7
20.6
41.8
Table 1: Quantitative Comparison. Our method outperforms recent 3D layout generation baselines across all evaluation metrics. The Ref. column gives how each method receives the reference scene and a dash marks methods that receive only the target floor plan. Best and second-best results are formatted in bold and underlined .
Figure 4: Qualitative Comparison. Our hierarchical framework preserves fine-grained spatial relationship across diverse room geometries while strictly respecting physical constraints.
Method
Physical Plausibility
Semantic Coherency
Overall
Fidelity
CF ↑
IB ↑
Pos. ↑
Rot. ↑
PSA ↑
iRecall ↑
(a) w/o Cluster Placement
94.5
94.6
47.4
47.9
48.8
51.1
(b) w/o Analogical Refinement
94.0
99.2
59.6
53.9
56.0
57.0
(c) w/o Physical Constraints
88.4
92.7
57.2
56.1
53.2
61.4
(d) w/o OpenShape
95.9
99.6
56.1
55.2
55.1
59.8
(e) w/o Concerto
95.3
99.8
52.5
51.1
51.9
54.6
Table 2: Ablation Study. Pipeline components (a–c) and conditioning signals (d–f) are evaluated. Best and second-best results are formatted in bold and underlined .
Figure 5: Downstream Applications. Our framework enables various applications such as (a) human and (b) camera trajectory transfer, (c) real-to-sim and (d) multiple reference composition.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Network Architecture. (a) Cluster placement network for predicting cluster poses conditioned on the target floor plan. (b) Object-level layout decoder for predicting object poses from the placed reference objects and target floor plan.
Rule
Labels
Distance
Angle
Adjacency
in_front_of , behind ,
gap <0.22
±45∘ quadrants
beside_left , beside_right
Facing
facing
≤0.6
±9∘ cone
Alignment
aligned_with
≤0.6
headings within 20∘
Wall
against_wall
<0.12 to boundary
N/A
Appendix
Table 3: Relation Extraction Rules. Distances are in normalized room coordinates; the adjacency gap is center distance minus the combined half-extents of both objects.
Reference cluster size
3
5
7
9
11
IB (before projection) ↑
93.7
93.9
91.7
88.6
87.9
iRecall ↑
56.7
62.9
64.3
64.2
64.2
Appendix
Table 4: Effect of Cluster Granularity. In boundary rate is measured before the projection step, and iRecall was measured for relational fidelity.
Figure 7: Cross-Scene Attention Visualization. Target object tokens selectively attend to their semantically corresponding reference furniture clusters, demonstrating robust semantic alignment under inventory and layout mismatches.
Figure 8: Same-Inventory Relocation. Paired reference and target room examples showing layout adaptation with identical object inventories across diverse room geometries.
Figure 9: Effect of Wall Clearance and Opening Terms. Omitting Lwall alters relative wall clearance, whereas omitting Lopen obstructs room access. The full model succeeds on both fronts.
Figure 10: Visualization of Ablation Study. Stage 1 placement preserves global semantic coherence, scene features provide essential conditioning, and Stage 3 refines object relationships.
Figure 11: Failure Scenarios. (a) Severe inventory mismatches between reference and target rooms lead to unnatural arrangements. (b) Highly constrained rooms cause conflicting spatial demands.
Figure 12: Additional Qualitative Comparisons. Extended visual comparisons between baseline methods and our framework across diverse 3D-Front ( Fu et al., 2021a ) scenes.
Dept. of Electrical and Computer Engineering, Seoul National University · Interdisciplinary Program in Artificial Intelligence and INMC, Seoul National University