Reference-based object compositing inserts or replaces an object using a background image, a reference image, and a 2D compositing mask. These inputs guide appearance and placement but leave the completed scene's geometry implicit, which can distort object structure or alter the surroundings. Our Depth-to-RGB (D2R) framework predicts composite depth for a scene not yet observed in the RGB inputs. It learns reference-conditioned corrections to a frozen depth estimator using encoder features of paired completed scenes as targets. The unchanged decoder maps the corrected representation to the intended scene's depth, which a separately trained renderer holds fixed during RGB synthesis. Under matched architecture and training, encoder-feature supervision reduces OOD Stage-1 AbsRel by 31.4% relative to decoded-depth supervision. We also introduce AnyInsertion++ with paired in-distribution and category-disjoint splits to evaluate generalization beyond compositing training categories. The complete D2R system leads 12 open-source and 3 closed-source baselines in estimator-derived geometry and photometric quality on both paired splits. On category-disjoint data, D2R reduces AbsRel by 43.7% and improves PSNR by 2.4 dB over the matched RGB baseline. Across three unpaired benchmarks, D2R leads both identity metrics and reduces mean CLIP reference cosine distance by 55% relative to the strongest baseline. Project page: https://shjo-april.github.io/Depth2RGB/
Figures & tables
Figure 1: Predicting Composite Depth Before RGB Synthesis. Depth-to-RGB (D2R; Ours) predicts composite depth (seventh column) and holds it fixed while generating RGB (eighth). D2R preserves reference appearance while matching the target shape and pose; the open- and closed-source baselines instead miss the target geometry, alter object identity, or modify surrounding content.
Method
Geometry Condition
Standard Inputs Only
Predicts Composite Depth
Fixed Geometry Before RGB
Insert-Anything [ 48 ] AAAI’26
Implicit in RGB
✓
✗
✗
Personalize-Anything [ 11 ] AAAI’26
Implicit in RGB
✓
✗
✗
SHINE [ 35 ] ICLR’26
Implicit in RGB
✓
✗
✗
HiFi-Inpaint [ 33 ] CVPR’26
Implicit in RGB
✓
✗
✗
MimicBrush [ 5 ] NeurIPS’24
Input Background Depth
✓
✗
✓
BIFRÖST [ 27 ] NeurIPS’24
Rule-Assembled Depth
✗
✗
✓
Table 1: Geometry for Reference-Based Object Compositing. Standard inputs are background and reference images plus a 2D compositing mask. Fixed geometry denotes depth maps or RGB proxies supplied before RGB generation.
Figure 2: Overview of Depth-to-RGB (D2R). Stage 1 (Sec. 3.1 ) predicts composite depth through reference-conditioned corrections to background features in a frozen depth estimator. Stage 2 (Sec. 3.2 ) synthesizes RGB with this depth held fixed throughout sampling. Training corrupts a separate canvas containing paired RGB, the reference, and target depth D . Depth denoising and RGB-depth consistency always target D , preventing errors in Din from becoming supervision.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Task and Inputs
Geometry Condition
Learned Components
Compose-and-Conquer [ 24 ] ICLR’24
Composable synthesis from text, foreground/background depth, and exemplar semantics
Supplied foreground/background depth maps
Local/global fusers and cloned diffusion blocks; original diffusion backbone frozen
3DIS [ 77 ] ICLR’25
Multi-instance generation from text and instance layouts
Scene depth generated before RGB
Fine-tuned depth generator and layout adapter; pretrained depth control and detail rendering without additional training
D2R (Ours)
Reference-based compositing from background, reference, and 2D compositing mask
Composite depth predicted before RGB
Depth-feature injector with frozen estimator; renderer LoRA
Appendix
Table A2: Geometry Inputs and Adaptation in Scene Synthesis.
Method
Conditioning and Task
Prediction Target
Learned Components
VGGT-World [ 51 ] ECCV’26
Video history for future geometry forecasting
Future geometry tokens, decoded into depth and point maps
Temporal flow transformer; VGGT encoder, remaining blocks, and 3D heads frozen
JointDiT [ 23 ] ICCV’25
Text, with optional RGB or depth for conditional generation
Joint or conditional RGB/depth latents
Depth-branch LoRA and cross-branch modules; pretrained FLUX backbone frozen
Modality Forcing [ 9 ] arXiv’26
Text, with optional RGB or depth for conditional generation
RGB latents and pixel-space depth
Post-trained DiT with modality-specific modules and noise levels; supports sparse depth supervision
D2R (Ours)
Background, reference, and 2D compositing mask for compositing
Completed-scene encoder features, decoded into composite depth
Reference-conditioned injector; depth encoder and decoder frozen, followed by renderer LoRA
Appendix
Table A3: Prediction Targets and Adaptation Across Geometry Models.
Split
Interior
Boundary Band
Background
ID
0.1081
0.2072
0.0522
OOD
0.1711
0.2422
0.0909
Appendix
Table B2: Regional AbsRel with Decoded-Depth Supervision.
Property
Original AnyInsertion [ 48 ]
AnyInsertion++ (Ours)
Paired Test Cases
38
287 ID and 139 OOD
Held-Out Object Tags
Not separated
21 tags, OOD split only
Training Pool Examined
77,745 before curation
52,918 kept, 24,827 removed
Ground-Truth Composite Depth
None
One per paired sample
Matched Comparator
None
Insert-Anything [ 48 ] retrained on ours
Appendix
Table C1: AnyInsertion++ Dataset Design.
Depth Predictor
Prediction
ID AbsRel ↓
OOD AbsRel ↓
Frozen Depth Anything V2 ViT-L [ 65 ] NeurIPS’24
Single Pass
0.1168
0.1506
Diffusion Depth (FLUX), No Fine-Tuning
28 Steps
0.1262
0.1792
Diffusion Depth (FLUX), Fine-Tuned
28 Steps
0.1067
0.1740
Encoder Feature Injection (D2R; Ours)
Single Pass
0.0462
0.0751
Appendix
Table C5: Composite Depth Prediction with Feature Correction and Diffusion.
Figure C1: Denoising Trajectories with Joint Prediction and Fixed Ground-Truth Depth. The upper branch predicts depth jointly with RGB; the lower D2R branch uses GT composite depth as a fixed condition. Tiles show one-step estimates from one example. Table 6 instead evaluates D2R using composite depth predicted by Stage 1.
Figure D1: Qualitative Results on the AnyInsertion++ In-Distribution (ID) Split.
Figure D2: Qualitative Results on the AnyInsertion++ Out-of-Distribution (OOD) Split.
Figure D3: Qualitative Results on TF-ICON [ 36 ] .
Figure D4: Qualitative Results on AIComposer [ 26 ] .
Figure D5: Qualitative Results on ComplexCompo [ 35 ] .