Reconstructing complete, scene-aligned 3D objects from casual images requires integrating sparse, uncertain observations and inferring surfaces hidden by occlusions. We present GATOR, a generative and agentic framework that recovers textured object assets and their scene-relative pose from one or more images. Our local modality mixer couples patch-aligned RGB, target-mask, and pointmap features before cross-view reasoning, preserving scene context while distinguishing the target from its surroundings. Text-guided semantic conditioning complements these spatial cues with category names and object captions through stage-specific adapters for structure, geometry, and appearance generation. The generated asset initializes a multimodal agent, providing instance-specific geometry and pose for targeted structural and texture refinement through an observation-guided edit-render-review loop. Across synthetic objects, cluttered tabletops, and indoor scenes, GATOR achieves strong geometric and appearance fidelity while recovering scene-relative pose from sparse observations. Time-budget comparisons and scene-level simulation further demonstrate the reconstruction efficiency and simulation readiness. Project page: https://research.nvidia.com/labs/lpr/gator/
Figures & tables
Fig. 1: Generative and agentic 3D object reconstruction from casual images. Left: Generative model inaccurately fills the strainer’s mesh bowl. GPT-6-Astra struggles with strainer ear position, and simplifies the handle and ear shapes. Our GATOR generates and refines the asset, preserving its shape and mesh. Right: GATOR reconstructs complete, textured, posed objects from real-world tabletop captures and room scans, which are composed using their predicted poses. Project page: https://research.nvidia.com/labs/lpr/gator/
Fig. 2: GATOR architecture. (A) Pose-aware generative reconstruction. Following TRELLIS.2 ( Xiang et al., 2026 ) , GATOR uses conditional flows for sparse structure, geometry, and appearance. RGB images and depth-derived pointmaps retain scene context, while masks identify the target. Tokens from all three modalities share a patch grid with common 2D position and Plücker-ray encodings. A local mixer exchanges information at each patch within a view (Secs. 3.2 and 3.3 ). Cross-view attention and SigLIP2 object-name conditioning guide structure generation. At its active coordinates, projected DINOv3 and NAF-upsampled features from target-masked images are averaged across valid views and injected into every geometry and appearance denoising block. Global image tokens and T5Gemma2-encoded captions supply cross-attention. Shape latents condition PBR material generation (Sec. 3.4 ). ❄ denotes frozen encoders. (B) Agentic refinement. Initialized with the complete, scene-aligned asset, an agent uses input images and structural priors to complete parts, repair topology, and refine texture coordinates and materials. Its edit-render-review loop compares original and candidate renders under matched cameras and lighting to retain or revert local edits. Within the time budget, it selects the best inspected asset while preserving reliable regions and scene-relative placement (Sec. 3.5 ).
Methods
Geometric
Appearance
CD ↓
NC ↑
F1 ↑
IoU 2D↑
PSNR ↑
SSIM ↑
LPIPS ↓
GPT-6-Astra
0.013
0.754
0.614
0.774
17.11
0.865
0.214
TRELLIS.2 †
0.014
0.751
0.628
0.747
15.73
0.855
0.231
ReconViaGen
0.014
0.767
0.579
0.764
14.39
0.825
0.228
Pixal3D-SV
0.027
0.681
0.433
0.629
14.22
0.844
0.287
Pixal3D-MV
0.011
0.783
0.701
0.798
17.18
0.871
0.203
Table 1: Four-view reconstruction on Toys4K. † Run with MultiDiffusion ( Bar-Tal et al., 2023 ) . Dark/light green mark best/second-best values (excluding the fully proprietary GPT-6-Astra baseline).
Methods
Joint (ADD-SB)
Geometry
Appearance
Mean ↓
@0.1 ↑
CD ↓
NC ↑
F1 ↑
IoU 2D↑
PSNR ↑
SSIM ↑
LPIPS ↓
GPT-6-Astra
0.027
100.00
0.016
0.799
0.773
0.791
17.16
0.886
0.160
SimFoundry
0.183
63.36
0.040
0.716
0.533
0.626
14.90
0.875
0.200
RecGen
0.057
85.50
0.021
0.791
0.666
0.757
16.00
0.885
0.178
ReconViaGen
—
—
0.027
0.750
0.608
0.730
14.48
0.858
0.181
ShapeR
0.057
87.02
0.028
0.709
0.606
0.700
—
—
—
Table 2: Real-world tabletop reconstruction on LM-O, HB, and HANDAL. ReconViaGen and Pixal3D lack scene placement, so joint ADD-SB is unavailable. ShapeR can only output untextured object meshes, so appearance scores are unavailable. Dark/light green mark best/second-best scores, excluding the fully proprietary GPT-6-Astra baseline. GATOR leads all metrics among the remaining methods, roughly halving Pixal3D’s chamfer distance (per-dataset results: Table 7 ).
Fig. 3: Real-world tabletop reconstruction on HB, LM-O, and HANDAL from four input views (two for RecGen); GPT-6 reconstructs from scratch. GATOR recovers complete shapes and detailed textures, such as the printed box and the whisk’s wires, where baselines truncate, blur, or fail.
Methods
Joint (ADD-SB)
Geometry
Mean ↓
@0.1 ↑
CD ↓
NC ↑
F1 ↑
IoU 2D↑
GPT-6-Astra
0.049
90.83
0.033
0.716
0.581
0.716
SimFoundry
0.125
72.49
0.104
0.664
0.426
0.614
RecGen
0.067
90.83
0.041
0.703
0.540
0.682
ReconViaGen
—
—
0.161
0.676
0.423
0.498
ShapeR
0.071
81.66
0.060
0.684
0.458
0.610
Table 3: Real-world indoor-scene reconstruction on ScanNet++. ReconViaGen and Pixal3D predict no pose placement and thus have no joint (ADD-SB) result. Dark/light green mark best/second-best values (excluding the fully proprietary GPT-6-Astra baseline).
Fig. 4: Reconstruction from real-world ScanNet++ images. Green outlines mark targets in a subset of the input views, and meshes are shown in comparable orientations; GPT-6 reconstructs from scratch. Our GATOR recovers fine structures, including individual utensils, fan blades, and the cart-mounted display.
Table 4: Ablation on 55 HANDAL objects with four input views. Each row adds one component to the row above. The modality mixer lowers ADD-SB, text conditioning improves all nine metrics, and agentic refinement further improves CD, NC, F1, PSNR, and LPIPS.
Fig. 5: Reconstruction quality versus time. (a) Mean CD/ D and LPIPS ( ↓ ) over valid outputs on sampled Toys4K objects. Time is the mean recorded full pipeline time, with an estimated 32 s of generation per object for GATOR, and excludes mesh export. The first GATOR point is generation without refinement (about 0.5 min), which already outperforms GPT-6 run with a 20-minute budget. (b) One example rendered from a shared camera. The first column shows our result without refinement, and the remaining headings denote agent budgets. GPT-6 times out within 2 min (sad faces) and remains coarse even at 20 min.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Scene source
Scenes
Epoch items
Unique assets
Share (%)
Reuse
S3
199,106
2,287,840
949,212
68.44
2.41 ×
3D-FRONT
10,506
63,624
9,454
1.90
6.73 ×
InternScenes
23,402
248,302
83,915
7.43
2.96 ×
SceneSmith
541
19,993
9,947
0.60
2.01 ×
SAGE
9,999
723,117
492,743
21.63
1.47 ×
Total
243,554
3,342,876
1,545,271
100.00
2.16 ×
Appendix
Table 6: Statistics of synthetic training corpus of the structure model. Epoch items are scene–object instances, share is their fraction per source, and reuse is the number of items per asset. Assets are counted within each source, so the total is not deduplicated across sources.
Fig. 6: S3 scene examples. Each adjacent image pair shows two viewpoints of one synthetic scene. Procedural layouts and varied capture conditions produce changes in occlusion, object scale, background, and lighting while retaining complete object supervision.
Fig. 7: Four-view reconstruction on Toys4K. Each example shows its four inputs at left and each method’s geometry (top) and appearance (bottom) in a shared novel view after shape alignment.
Methods
Joint (ADD-SB)
Geometry
Appearance
Mean ↓
@0.1 ↑
CD ↓
NC ↑
F1 ↑
IoU 2D↑
PSNR ↑
SSIM ↑
LPIPS ↓
LM-O: sensor depth
GPT-6-Astra
0.045
100.00
0.031
0.777
0.524
0.789
15.92
0.899
0.243
SimFoundry
0.375
62.50
0.043
0.783
0.460
0.748
15.50
0.882
0.252
RecGen
0.043
100.00
0.027
0.796
0.542
0.811
16.95
0.911
0.225
ReconViaGen
—
—
0.049
0.723
0.413
0.728
12.58
0.837
0.256
Appendix
Table 7: Real-world tabletop reconstruction per dataset. The generative model alone (GATOR w/o agent) already outperforms every baseline on every geometry metric, and refinement improves PSNR and LPIPS on every dataset. Dashes mark unavailable scores, including appearance scores for ShapeR, which can only output untextured object meshes. Dark/light green mark best/second-best values per dataset (excluding the fully proprietary GPT-6-Astra baseline).
Fig. 8: Extended LM-O and HB comparisons. Each object shows two of its four input views and each method’s geometry (top) and appearance (bottom) from a shared camera. ShapeR can only output untextured object meshes (red crosses).
Views
Methods
Joint (ADD-SB)
Geometry
Appearance
Mean ↓
@0.1 ↑
CD ↓
NC ↑
F1 ↑
IoU 2D↑
PSNR ↑
SSIM ↑
LPIPS ↓
2
RecGen
0.041
98.03
0.021
0.796
0.643
0.845
15.16
0.867
0.203
Pixal3D
—
—
0.027
0.757
0.648
0.840
15.32
0.865
0.209
GATOR (Ours)
0.015
98.03
0.011
0.865
0.863
0.897
15.75
0.866
0.183
4
Pixal3D
—
—
0.021
0.799
0.745
0.875
16.29
0.872
0.190
GATOR (Ours)
0.014
98.03
0.010
0.874
0.891
0.902
15.86
0.866
0.179
Appendix
Table 8: Multiview tabletop reconstruction on LM-O and HB. Unlike Table 7 , each target is evaluated with two different view selections, recall counts every attempt, and GATOR is reported without agentic refinement, so its four-view scores differ between the two tables. GATOR leads every joint and geometry metric at every view count, and its distance errors decrease as views are added. Pixal3D predicts no pose placement and thus has no joint (ADD-SB) result. Bold marks the best score.
Fig. 9: Extended ScanNet++ comparisons. Each object shows up to two input views with the target outlined in green and each method’s geometry (top) and appearance (bottom) from a shared camera. ShapeR can only output untextured object meshes (red crosses). GT is the observed partial scan.
Fig. 10: Agentic reconstruction refinement. Before-and-after examples illustrate six refinement behaviors. Up to four input views are shown, with real-world targets outlined in green. Boxes and matched zooms highlight reconstruction details. Each pair shares the same camera, diffuse shading, and display scale.
Fig. 11: Scene-level rigid-body simulation. All ten input views (top two rows) and simulated resting states of a ScanNet++ meeting room, shown from overview (third row) and side (bottom) cameras. Three chairs that lean or topple before refinement become able to stand upright steadily after refinement.
Reconstructing complete 3D object assets from monocular or sparse multi-view observations remains challenging. Generative 3D foundation models can complete object geometry beyond the observed views, but their predictions may not faithfully reproduce the observed geometry, appearance, or pose. We introduce GenIA, a framework for test-time input-aligned generation that grounds SAM3D's generative prior in geometric and photometric observations without retraining the foundation model. We improve object pose by deriving translation and scale from geometry while retaining the learned rotation prior, and align appearance through visibility-biased attention, cross-observation fusion, and differentiable rendering guidance during denoising. An optional post-denoising refinement further adapts the appearance latent, lightweight decoder adapters, and object placement to the observations. Our framework also supports externally supplied geometry; when given temporal shapes of dynamic objects, it recovers a shared, input-aligned canonical appearance and stable world-space placement. Across synthetic and real benchmarks, GenIA improves pose prediction and object reconstruction from monocular, multi-view, and dynamic inputs, outperforming recent optimization-based, per-frame image-to-3D, and video-to-4D methods. Our project page is available at https://facebookresearch.github.io/GenIA.
Accurately reconstructing complex full multi-object scenes from sparse observations remains a core challenge in computer vision and a key step toward scalable and reliable simulation for robotics. In this work, we introduce RecGen, a generative framework for probabilistic joint estimation of object and part shapes, as well as their pose under occlusion and partial visibility from one or multiple RGB-D images. By leveraging compositional synthetic scene generation and strong 3D shape priors, RecGen generalizes across diverse object types and real-world environments. RecGen achieves state-of-the-art performance on complex, heavily occluded datasets, robustly handling severe occlusions, symmetric objects, object parts, and intricate geometry and texture. Despite using nearly 80% fewer training meshes than the previous state of the art SAM3D, RecGen outperforms it by 30.1% in geometric shape quality, 9.1% in texture reconstruction, and 33.9% in pose estimation.
We propose Tetris3D, a generative framework for single-image 3D scene reconstruction that recovers objects which are physically and geometrically coherent as a scene. Existing methods often generate objects independently or couple them implicitly, providing limited guidance for ensuring fine-grained spatial compatibility between neighboring objects that interact with one another. To address this, we explicitly condition the generation of each object on the geometry of surrounding objects and their physical relationships, guiding its shape and pose to remain geometrically and physically plausible within the scene. Moreover, we introduce ComOb, a physics simulation-based dataset of 1.2M scenes featuring physical interactions across diverse object categories, with per-object meshes and pairwise physical relation annotations. Comprehensive experiments on synthetic and realworld scenes show that Tetris3D recovers coherent object shapes and poses even when interacting regions are occluded, and achieves state-of-the-art performance in both generation quality and physical stability.