Reconstructing complete 3D object assets from monocular or sparse multi-view observations remains challenging. Generative 3D foundation models can complete object geometry beyond the observed views, but their predictions may not faithfully reproduce the observed geometry, appearance, or pose. We introduce GenIA, a framework for test-time input-aligned generation that grounds SAM3D's generative prior in geometric and photometric observations without retraining the foundation model. We improve object pose by deriving translation and scale from geometry while retaining the learned rotation prior, and align appearance through visibility-biased attention, cross-observation fusion, and differentiable rendering guidance during denoising. An optional post-denoising refinement further adapts the appearance latent, lightweight decoder adapters, and object placement to the observations. Our framework also supports externally supplied geometry; when given temporal shapes of dynamic objects, it recovers a shared, input-aligned canonical appearance and stable world-space placement. Across synthetic and real benchmarks, GenIA improves pose prediction and object reconstruction from monocular, multi-view, and dynamic inputs, outperforming recent optimization-based, per-frame image-to-3D, and video-to-4D methods. Our project page is available at https://facebookresearch.github.io/GenIA.
Figures & tables
Figure 2 : Our pipeline. Left to right, top to bottom, we obtain the coarse shape from SAM3D’s first pass or ( optionally ) inject an external one into SAM3D’s shape tokens ( Sec. 4.1 ), denoise each object’s per-frame pose over that fixed shape ( Sec. 4.2 ), predict observation-aligned appearance as SLAT features ( Sec. 4.3 ), and ( optionally ) do test-time refinement of the reconstruction against the input views ( Sec. 4.4 ). Dotted arrows denote gradient flow. We illustrate the single-view setting; multi-view and dynamic extensions are described in the text.
Figure 3 : Visibility on the object’s coarse shape. Left: camera view; orbiting reveals visible and occluded voxels.
Figure 4 : Qualitative comparison. 4(a) GSO-30 and CO3D reconstructed from one input view and from multi-view inputs ( 5 views for GSO-30, 4 for CO3D). 4(b) ActionBench and DAVIS reconstructed from 16-frame sequences; for DAVIS, we show the first and last reconstructed frames. Methods are described in Sec. 5 . DAVIS provides no ground-truth test views, so its test subrow shows synthesized novel views at the same timestamps. Each render is badged with per-scene PSNR ( ↑ ) on train views and CLIP-I ( ↑ ) on test views.
Table 1 : Full benchmarks. Per-dataset image quality metrics across methods and, where applicable, number of input views (#V). “–” denotes metrics not reported by the corresponding method. Baselines, datasets, metrics and geometric backbones used by our method are detailed in Sec. 5 . TRELLISv2 and Pixal3D predict lighting-disentangled materials and are therefore evaluated qualitatively only. Best and second best per (#V, dataset, metric) group are shown in bold and underlined, respectively.
Figure 5 : GenIA with TTR completes the unseen surface with colors consistent with the input. A real scene rendered from the side opposite the input view. SAM3D completes the object with a smooth, synthetic appearance driven by its generative prior. TripoSplat and CUPID closely match the observed side but degrade on unseen surfaces, with washed-out colors or appearance drift.
Figure 6 : Our method adds a small overhead over SAM3D. Median runtime versus mean CLIP-I, with runtime given as a multiple of SAM3D’s, which takes 25 s on our hardware (a single A100-40 GB). The two static panels are scored on CO3D; the dynamic panel is scored on DAVIS. Without TTR, rendering guidance is the main overhead: runtime is approximately affine in view count, with per-view cost comparable to multi-view baselines. TTR adds a largely view-independent cost, while improving CLIP-I. On dynamic sequences, GenIA costs 13× SAM3D’s single-image runtime including ActionMesh ( 12× without TTR), versus 160× for Lift4D and 250× for HiMoR, and with TTR achieves the highest CLIP-I.
Figure 7 : Our appearance prediction better aligns with the input observations. Incremental appearance ablation on GSO-30, following Sec. 5.2 . Attention bias injects pose information into SAM3D’s pose-blind appearance predictor, helping resolve ambiguous appearance assignment on symmetric coarse geometry. Rendering guidance further improves color and detail matching, while optional TTR provides additional gains, including on held-out views. We show one train and one held-out test view, badged with PSNR and co-PSNR ( ↑ ).
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Table 2: Pose ablations across datasets. Four-step cumulative ablation, with each row adding one component of Sec. 4.2 . We evaluate dynamic scenes (DAVIS and ActionBench with ActionMesh geometry) and static GSO-30 single- and multi-view settings with injected ground-truth geometry. Dynamic settings report reprojection error and relative jitter against the CoTracker3 pseudo-reference ( Sec. A.11 ), together with masked depth L1 . GSO-30 additionally measures placement of the injected coarse shape against ground truth using Chamfer distance. Jit. is relative: 1.0 matches the reference motion, while larger values indicate greater jitter. Rotation averaging is inert on static settings by construction, since it averages velocities across frames. ICP registration is most beneficial for static multi-view alignment, where it halves Chamfer distance; on DAVIS, it improves depth at the cost of appearance quality.
Table 3 : Appearance ablations. Per-dataset novel-view synthesis across the component ladder: the columns toggle Attn (visibility attention bias), RG (rendering guidance), and TTR (test-time refinement); ✓ / × mark a component on/off and “–” a component that does not apply (or a metric not reported). Held-out views report co-PSNR where the dataset has co-visibility masks, otherwise only LPIPS and CLIP-I. All GSO-30 and ActionBench variants share ground-truth shape and pose. CO3D uses SAM3D- and MV-SAM3D-predicted shape and pose. DAVIS uses the ActionMesh backbone. Metrics are averaged across scenes and views (GSO-30, CO3D), views and time (ActionBench), or time (DAVIS). Best and second best per (#V, dataset, metric) group are shown in bold and underlined, respectively.
Figure 8 : Qualitative appearance ablation on GSO-30 with ground-truth geometry. Every column is our method on the GT-voxelized mesh (shared GT shape and pose), so geometry is fixed and only appearance differs. Cumulative ladder, where each column adds one component to the column on its left. The label names only what that step adds: Base (SAM3D appearance prediction + observation fusion) → +Attn (visibility attention bias) → +RG (rendering guidance) → +TTR (test-time refinement). Each scene spans two subrows: rendered from the first input ( train ) view and from a held-out novel ( test ) view, each against its GT. Each render is badged (bottom-right) with its per-scene appearance score: PSNR ( ↑ ) on the train views and co-visible PSNR (co-PSNR, ↑ ; Sec. 5 ) on the test views.
Figure 9 : Qualitative appearance ablation on ActionBench with ground-truth geometry. Every column is our method on the per-frame GT mesh (shared GT shape and pose), so only appearance differs. Cumulative ladder, where each column adds one component to the column on its left. The label names only what that step adds: Base (SAM3D appearance prediction + observation fusion) → +Attn (visibility attention bias) → +RG (rendering guidance) → +TTR (test-time refinement). Each scene spans two subrows: rendered from the input ( train ) view and from the held-out ( test ) view, each against its GT. Each render is badged (bottom-right) with its per-scene appearance score: PSNR ( ↑ ) on the train views and co-visible PSNR (co-PSNR, ↑ ; Sec. 5 ) on the test views.
Figure 10 : Qualitative appearance ablation on CO3D with SAM3D geometry. Every column is our method on predicted shape and pose (CO3D has no ground-truth mesh); SAM3D geometry from 1 input view and MV-SAM3D geometry from 4 input views. Only appearance differs. Cumulative ladder, where each column adds one component to the column on its left. The label names only what that step adds: Base (SAM3D appearance prediction + observation fusion) → +Attn (visibility attention bias) → +RG (rendering guidance) → +TTR (test-time refinement). Each scene spans two subrows: the reconstruction rendered from the first input ( train ) view and from a held-out target ( test ) view, each against its GT. Each render is badged (bottom-right) with its per-scene appearance score: PSNR ( ↑ ) on the train views and CLIP-I ( ↑ ) on the test views.
Figure 11 : Qualitative appearance ablation on DAVIS with ActionMesh geometry. Every column is GenIA on the same ActionMesh-predicted temporal coarse shape, so only appearance differs. Cumulative ladder, where each column adds one component to the column on its left. The label names only what that step adds: Base (SAM3D appearance prediction + observation fusion) → +Attn (visibility attention bias) → +RG (rendering guidance) → +TTR (test-time refinement). Each scene spans two subrows: the reconstruction rendered from the input ( train ) view against its GT (the input frame foreground-composited on white via the DAVIS annotation mask), and a synthesized novel view at the same timestamp ( test ). Each render is badged with per-scene PSNR ( ↑ ) on train views and CLIP-I ( ↑ ) on test views.
Figure 12 : Qualitative comparison on GSO-30. Each render is badged (bottom-right) with its per-scene PSNR ( ↑ ) on the train views and CLIP-I ( ↑ ) on the test views.
Figure 13 : Qualitative comparison on ActionBench. Each render is badged (bottom-right) with its per-scene PSNR ( ↑ ) on the train views and CLIP-I ( ↑ ) on the test views.
Figure 14 : Qualitative comparison on CO3D. Novel-view synthesis on 4 held-out CO3D (LaRa) scenes, reconstructed from 1, 2 and 4 input views. We compare each setting’s baselines (SAM3D, TripoSplat, TRELLISv2, Pixal3D, RecGen, and CUPID at 1 view; MV-SAM3D and RecGen at 2 views; MV-SAM3D, Depth-3DGS, STREAM3D, and ReconViaGen at 4 views) against our method on that setting’s geometry backbone (SAM3D at 1 view, MV-SAM3D at 2 and 4 views). RecGen’s released checkpoint takes at most two views, so we show it at its maximum, in the 2-view group. TRELLISv2 and Pixal3D are qualitative-only baselines: they predict mesh materials without environment lighting, so their renders show unshaded base color ( Sec. 5 ). Each scene spans two subrows: the reconstruction rendered from the first input ( train ) view and from the held-out target ( test ) view, each against its GT. Each render is badged (bottom-right) with its per-scene PSNR ( ↑ ) on the train views and CLIP-I ( ↑ ) on the test views.
Figure 15 : Qualitative comparison on DAVIS. Each scene spans two subrows. The train subrow re-renders the input viewpoint, against the input frame as GT. DAVIS is monocular and has no second camera, so the test subrow is instead a synthesized novel view of the same timestamp. Each reconstruction is badged (bottom-right) with its per-scene PSNR ( ↑ ).
Figure 16 : Failure cases. Dynamic geometry currently relies on an external predictor (ActionMesh) rather than being directly grounded in the observations. Rotation stays prior-driven, corrected only by noise-sensitive ICP registration. Both failure modes surface on dynamic scenes below: incorrect world-space placement from registration, non-rigid deformation errors inherited from ActionMesh, or a combination of the two, can strongly impact the final reconstruction quality. Each render is badged with per-scene PSNR ( ↑ ) on train views and CLIP-I ( ↑ ) on test views (DAVIS shows a synthesized novel view in place of a held-out test view; its GT test tile is left blank).
Figure 17 : Per-frame to canonical voxel correspondence on ActionBench (scene 000-003 ). Each frame’s mesh is voxelized in its own bounding box (its native per-frame grid); every voxel is matched to a canonical voxel through the shared fixed topology ( Sec. A.3 ) and colored by that canonical voxel’s position, so a correct correspondence keeps each body part a constant color throughout the sequence. In this example, columns show the canonical frame c=5 and frames 1,6,11,16 ; rows show three orbit viewpoints.
Figure 18 : Simulated effect of passive-stream compensation for visibility-biased attention. We visualize four simulated DINOv2 37×37 streams. Visibility bias ( α>0 ) boosts projected patches in the active RGB streams (CroppedImg, FullImg), which in turn suppresses the passive mask streams (CroppedMask, FullMask) through softmax normalization. Without compensation ( off ), passive-stream attention decreases; the closed-form approximation ( approx ) nearly restores its unbiased level and closely matches the post-softmax oracle ( exact ).
Generating pose-aligned 3D objects is challenging due to the spatial mismatches and transformation ambiguities inherent in decoupled canonical-then-rotate paradigms. To this end, we introduce Pose-Aware Diffusion (PAD), a novel end-to-end diffusion framework that synthesizes 3D geometry directly within the observation space. By unprojecting monocular depth into a partial point cloud and explicitly injecting it as a 3D geometric anchor, PAD abandons canonical assumptions to enforce rigorous spatial supervision. This native generation intrinsically resolves pose ambiguity, producing high-fidelity pose-aligned assets. Extensive experiments demonstrate that PAD achieves superior geometric alignment and image-to-3D correspondence compared to state-of-the-art methods. Additionally, PAD naturally extends to compositional 3D scene reconstruction via a simple union of independently generated objects, highlighting its robust ability to preserve precise spatial layouts.
Zihan Zhou, Luxi Chen, Jingzhi Zhou +4
Gaoling School of AI, Renmin University of China · VCIP, School of Computer Science, Nankai University · THU-Bosch MLCenter, Tsinghua University +1
Reconstructing complete, scene-aligned 3D objects from casual images requires integrating sparse, uncertain observations and inferring surfaces hidden by occlusions. We present GATOR, a generative and agentic framework that recovers textured object assets and their scene-relative pose from one or more images. Our local modality mixer couples patch-aligned RGB, target-mask, and pointmap features before cross-view reasoning, preserving scene context while distinguishing the target from its surroundings. Text-guided semantic conditioning complements these spatial cues with category names and object captions through stage-specific adapters for structure, geometry, and appearance generation. The generated asset initializes a multimodal agent, providing instance-specific geometry and pose for targeted structural and texture refinement through an observation-guided edit-render-review loop. Across synthetic objects, cluttered tabletops, and indoor scenes, GATOR achieves strong geometric and appearance fidelity while recovering scene-relative pose from sparse observations. Time-budget comparisons and scene-level simulation further demonstrate the reconstruction efficiency and simulation readiness. Project page: https://research.nvidia.com/labs/lpr/gator/
Recovering complete 3D representations of objects from few casual image captures remains a significant challenge. Recent 3D generative models, particularly those based on Flow-Matching (FM), can synthesize high-quality textured assets; however, they often suffer from ''synthetic bias'' where learned priors override observational evidence, alongside a lack of alignment with the observed instance. Conversely, optimization-based methods like 3D Gaussian Splatting (3DGS) provide high fidelity on visible surfaces but fail to reason about unobserved geometry. In this paper, we present FlowObject, a framework that reformulates sparse-view 3D reconstruction as a training-free, guided inverse problem. Our approach applies a dual-space guidance strategy to steer the Ordinary Differential Equation (ODE) trajectory of a flow-matching model, enabling the completion of unseen regions through learned generative priors while enforcing strict consistency with real-world observations. By integrating a 3DGS refinement stage, FlowObject further bridges the gap between ''synthetic-looking'' generative outputs and photorealistic reconstructions. Comprehensive benchmarks on synthetic and real-world datasets demonstrate that current state-of-the-art methods often struggle to achieve geometric completeness and observational consistency simultaneously, especially under severe occlusions. In contrast, our method significantly outperforms state-of-the-art generative models and optimization-based frameworks in both geometric completeness and view-dependent appearance fidelity.
Yuchen Rao, Xuqian Ren, Yinyu Nie +4
Graz University of Technology Austria · Tampere University Finland · Technical University of Munich Germany +3