Reference-based object compositing inserts or replaces an object using a background image, a reference image, and a 2D compositing mask. These inputs guide appearance and placement but leave the completed scene's geometry implicit, which can distort object structure or alter the surroundings. Our Depth-to-RGB (D2R) framework predicts composite depth for a scene not yet observed in the RGB inputs. It learns reference-conditioned corrections to a frozen depth estimator using encoder features of paired completed scenes as targets. The unchanged decoder maps the corrected representation to the intended scene's depth, which a separately trained renderer holds fixed during RGB synthesis. Under matched architecture and training, encoder-feature supervision reduces OOD Stage-1 AbsRel by 31.4% relative to decoded-depth supervision. We also introduce AnyInsertion++ with paired in-distribution and category-disjoint splits to evaluate generalization beyond compositing training categories. The complete D2R system leads 12 open-source and 3 closed-source baselines in estimator-derived geometry and photometric quality on both paired splits. On category-disjoint data, D2R reduces AbsRel by 43.7% and improves PSNR by 2.4 dB over the matched RGB baseline. Across three unpaired benchmarks, D2R leads both identity metrics and reduces mean CLIP reference cosine distance by 55% relative to the strongest baseline. Project page: https://shjo-april.github.io/Depth2RGB/
Figures & tables
Figure 1: Predicting Composite Depth Before RGB Synthesis. Depth-to-RGB (D2R; Ours) predicts composite depth (seventh column) and holds it fixed while generating RGB (eighth). D2R preserves reference appearance while matching the target shape and pose; the open- and closed-source baselines instead miss the target geometry, alter object identity, or modify surrounding content.
Method
Geometry Condition
Standard Inputs Only
Predicts Composite Depth
Fixed Geometry Before RGB
Insert-Anything [ 48 ] AAAI’26
Implicit in RGB
✓
✗
✗
Personalize-Anything [ 11 ] AAAI’26
Implicit in RGB
✓
✗
✗
SHINE [ 35 ] ICLR’26
Implicit in RGB
✓
✗
✗
HiFi-Inpaint [ 33 ] CVPR’26
Implicit in RGB
✓
✗
✗
MimicBrush [ 5 ] NeurIPS’24
Input Background Depth
✓
✗
✓
BIFRÖST [ 27 ] NeurIPS’24
Rule-Assembled Depth
✗
✗
✓
Table 1: Geometry for Reference-Based Object Compositing. Standard inputs are background and reference images plus a 2D compositing mask. Fixed geometry denotes depth maps or RGB proxies supplied before RGB generation.
Figure 2: Overview of Depth-to-RGB (D2R). Stage 1 (Sec. 3.1 ) predicts composite depth through reference-conditioned corrections to background features in a frozen depth estimator. Stage 2 (Sec. 3.2 ) synthesizes RGB with this depth held fixed throughout sampling. Training corrupts a separate canvas containing paired RGB, the reference, and target depth D . Depth denoising and RGB-depth consistency always target D , preventing errors in Din from becoming supervision.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Task and Inputs
Geometry Condition
Learned Components
Compose-and-Conquer [ 24 ] ICLR’24
Composable synthesis from text, foreground/background depth, and exemplar semantics
Supplied foreground/background depth maps
Local/global fusers and cloned diffusion blocks; original diffusion backbone frozen
3DIS [ 77 ] ICLR’25
Multi-instance generation from text and instance layouts
Scene depth generated before RGB
Fine-tuned depth generator and layout adapter; pretrained depth control and detail rendering without additional training
D2R (Ours)
Reference-based compositing from background, reference, and 2D compositing mask
Composite depth predicted before RGB
Depth-feature injector with frozen estimator; renderer LoRA
Appendix
Table A2: Geometry Inputs and Adaptation in Scene Synthesis.
Method
Conditioning and Task
Prediction Target
Learned Components
VGGT-World [ 51 ] ECCV’26
Video history for future geometry forecasting
Future geometry tokens, decoded into depth and point maps
Temporal flow transformer; VGGT encoder, remaining blocks, and 3D heads frozen
JointDiT [ 23 ] ICCV’25
Text, with optional RGB or depth for conditional generation
Joint or conditional RGB/depth latents
Depth-branch LoRA and cross-branch modules; pretrained FLUX backbone frozen
Modality Forcing [ 9 ] arXiv’26
Text, with optional RGB or depth for conditional generation
RGB latents and pixel-space depth
Post-trained DiT with modality-specific modules and noise levels; supports sparse depth supervision
D2R (Ours)
Background, reference, and 2D compositing mask for compositing
Completed-scene encoder features, decoded into composite depth
Reference-conditioned injector; depth encoder and decoder frozen, followed by renderer LoRA
Appendix
Table A3: Prediction Targets and Adaptation Across Geometry Models.
Split
Interior
Boundary Band
Background
ID
0.1081
0.2072
0.0522
OOD
0.1711
0.2422
0.0909
Appendix
Table B2: Regional AbsRel with Decoded-Depth Supervision.
Property
Original AnyInsertion [ 48 ]
AnyInsertion++ (Ours)
Paired Test Cases
38
287 ID and 139 OOD
Held-Out Object Tags
Not separated
21 tags, OOD split only
Training Pool Examined
77,745 before curation
52,918 kept, 24,827 removed
Ground-Truth Composite Depth
None
One per paired sample
Matched Comparator
None
Insert-Anything [ 48 ] retrained on ours
Appendix
Table C1: AnyInsertion++ Dataset Design.
Depth Predictor
Prediction
ID AbsRel ↓
OOD AbsRel ↓
Frozen Depth Anything V2 ViT-L [ 65 ] NeurIPS’24
Single Pass
0.1168
0.1506
Diffusion Depth (FLUX), No Fine-Tuning
28 Steps
0.1262
0.1792
Diffusion Depth (FLUX), Fine-Tuned
28 Steps
0.1067
0.1740
Encoder Feature Injection (D2R; Ours)
Single Pass
0.0462
0.0751
Appendix
Table C5: Composite Depth Prediction with Feature Correction and Diffusion.
Figure C1: Denoising Trajectories with Joint Prediction and Fixed Ground-Truth Depth. The upper branch predicts depth jointly with RGB; the lower D2R branch uses GT composite depth as a fixed condition. Tiles show one-step estimates from one example. Table 6 instead evaluates D2R using composite depth predicted by Stage 1.
Figure D1: Qualitative Results on the AnyInsertion++ In-Distribution (ID) Split.
Figure D2: Qualitative Results on the AnyInsertion++ Out-of-Distribution (OOD) Split.
Figure D3: Qualitative Results on TF-ICON [ 36 ] .
Figure D4: Qualitative Results on AIComposer [ 26 ] .
Figure D5: Qualitative Results on ComplexCompo [ 35 ] .
Depth can resolve appearance ambiguity in RGB-D salient object detection (SOD), yet sensor depth is not uniformly reliable. Missing regions, blurred boundaries, and structural artifacts can propagate through multimodal fusion and make an RGB-D detector less accurate than its RGB-only counterpart. Existing quality-aware approaches regulate observed depth but remain dependent on the same potentially defective modality. We propose \method, a reliability-aware geometry distillation framework developed for RGB-D SOD benchmarks without using dataset-provided depth during training or inference. A frozen Depth Anything V2 model serves only as a training-time teacher, transferring dense relative geometry, hierarchical spatial attention, and boundary structure to a compact edge-aware geometry branch. Pooled bidirectional interaction aligns geometry with appearance, and a pixel-wise reliability estimator selectively injects geometry that is compatible with the current RGB representation. The teacher is removed after training, leaving an RGB-only inference network. Trained on 2,985 RGB-mask pairs, \method{} achieves the best or tied-best result in 26 of 36 metric-dataset comparisons against ten recent RGB-D SOD methods, including a 13.4% relative MAE reduction on ReDWeb-S. When retrained on DUTS-TR, it also improves the strongest prior F-measure by 4.2% on PASCAL-S, showing that the distilled geometry transfers beyond a particular sensor or dataset domain. Code will be released upon publication.
Xuehao Wang, Jiaxin Hua, Runmei Li +4
University of International Business and Economics · State Key Laboratory of Virtual Reality Technology and Systems · Southwest Jiaotong University +2
Object insertion aims to seamlessly composite a reference object into a specified region of a background image. Recent diffusion-based methods achieve high visual quality but formulate insertion as a simple 2D inpainting task, providing no explicit control over the object's 3D pose and limiting their practical applicability. We propose DIRECT (Decomposed Injection for Reference Composition and Target-integration), a novel framework that integrates interactive pose manipulation with high-fidelity 2D image synthesis to enable pose-controllable object insertion. Our method decomposes the insertion conditions into three complementary components: appearance guidance capturing visual details from the reference object, geometry guidance derived from the user-adjusted 3D proxy, and context guidance from the target background. By injecting them through separate pathways, DIRECT avoids feature entanglement and simultaneously preserves reference appearance, follows the user-specified pose, and adapts the object to the target scene. We also introduce an automated data construction pipeline to improve the diversity and quality of training data. Experiments show that DIRECT outperforms previous methods in both geometric controllability and visual quality.
Jingbo Gong, Yikai Wang, Yushi Lan +6
VCIP, School of Computer Science, Nankai University · Zhongguancun Academy · S-Lab, Nanyang Technological University +2
Recent advances in robot imitation learning have produced visuomotor policies that predict actions directly from visual observations. Yet visually similar scenes can require different actions as target position, object height, or contact geometry changes. Pretrained RGB features may map these geometrically distinct states to similar policy inputs, while simply adding depth requires the policy to learn RGB-depth correspondence from the same limited demonstrations used to learn control. We introduce StereoPatch, a patch-aligned RGB-depth representation that binds registered metric geometry directly to the RGB patches used for action prediction. On a shared 2-D patch grid, asymmetric cross-attention incorporates depth information into the corresponding RGB features before action decoding. The resulting StereoPatch Tokens provide a geometry-aware visual representation that can condition general visuomotor policies without changing their underlying learning objectives. Across six real-robot tasks, StereoPatch achieves higher closed-loop success than appearance-only, geometry-only, raw RGB-D, and late-fusion baselines. Additional experiments across three simulation suites evaluate compatibility across visuomotor policy architectures, spatial generalization, and operating limits. Results suggest that resolving control-relevant geometric ambiguity benefits from aligning depth directly with the visual features used for action prediction, rather than supplying it as an independent modality. Project page: https://aus.bot/research/stereopatch/.
Yanan Zhou, Zhaoyan Qian, James Zhao +1
School of Computer Science, The University of Sydney, Australia. · Australian Centre for Robotics, The University of Sydney, Australia.