Decompositional scene reconstruction aims to reconstruct high-quality objects and background, yet existing methods still struggle with the level of quality under heavy occlusions. While generative priors offer a potential solution, 2D image-based priors often suffer from multi-view inconsistency due to a lack of 3D awareness. Conversely, 3D-native priors provide stronger structural inductive biases but frequently lead to spatial drift and misalignment within complex scenes. To address these issues, we propose DecomVoxel, formulating object completion as a guided in-situ denoising optimization that bridges 3D-native priors with neural scene reconstruction. Our framework introduces a reformulated epsilon-based distillation loss to ensure stable latent refinement, alongside adaptive spatial guidance that utilizes occupied and vacant anchors with temporal annealing to suppress generative hallucinations and mitigate spatial drift. Experiments on Replica and ScanNet++ show that DecomVoxel significantly outperforms state-of-the-art methods while faithfully preserving the original spatial layout, structural fidelity, and style-consistent texture. Our method pushes the boundary of decompositional reconstruction by delivering high-quality textured meshes with clean topology, geometry, and appearance, providing a robust solution for the decompositional reconstruction of complex real-world scenes. Code is available at https://github.com/DecomVoxel/DecomVoxel.
Figures & tables
Figure 1 . We propose DecomVoxel , a framework that integrates guided 3D-native priors to enhance decompositional scene reconstruction. By formulating object completion as guided in-situ optimization, our method achieves high-quality topology, geometry, and appearance for both individual objects and backgrounds while faithfully preserving the original spatial layout.
Figure 2 . Overview of DecomVoxel . Our framework bridges 3D generative priors with neural reconstruction through a guided in-situ denoising optimization. We segment objects from the initial scene and then sequentially optimize their geometry and appearance under adaptive spatial guidance. This process recovers missing structures and consistent textures while preserving the original scene layout, producing high-fidelity topology, geometry, and appearance.
Figure 3 . Visualization of intermediate results. The left panel illustrates our spatial anchors: the occupied anchor ΩO (purple) preserves verified structures, while the vacant anchor ΩV (red) suppresses hallucinations in prohibited zones. The green regions denote the free space for structural synthesis. The right panel displays the progressive geometric completion over iterations. Guided by these deterministic anchors, the generative process restores missing regions while maintaining the original scene layout.
Figure 4 . Qualitative comparison. Our method reconstructs complete geometry and consistent appearance while faithfully preserving the original scene layout, whereas baselines exhibit substantial drift in position, rotation, and scale.
Dataset
Method
Reconstruction
Rendering
CD ↓
F-score ↑
NC ↑
PSNR ↑
SSIM ↑
LPIPS ↓
Replica
SimRecon
17.55
40.61
63.25
18.04
0.858
0.152
ReconViaGen
15.49
51.02
68.09
18.58
0.861
0.147
CUPID
16.29
43.95
64.80
16.70
0.809
0.180
SAM3D
10.59
45.46
66.33
18.89
0.847
0.143
MV-SAM3D
12.87
40.46
64.19
18.83
0.848
0.142
Table 1. Quantitative comparison. Our method consistently outperforms state-of-the-art baselines across both reconstruction and rendering metrics. Top-3 results are highlighted as the first , second and third .
Method
Obj Geo. ↑
Obj App. ↑
Layout ↑
SimRecon
2.12±1.30
1.86±1.25
2.01±1.12
SAM3D
2.37±1.40
2.21±1.41
2.26±1.22
MV-SAM3D
2.66±1.59
2.52±1.60
2.68±1.01
ShapeR
2.19±1.36
-
2.95±1.20
Ours
3.81±1.25
3.78±1.33
4.08±0.95
Table 2 . User Study. Based on 44 expert questionnaires, we assess the scores (rated 1–5) of object completeness and consistency in geometry and appearance, along with spatial-layout fidelity relative to the original scene. Our method significantly outperforms all baselines across all metrics, particularly in preserving global structural alignment.
GP
G-SG
AP
A-SG
Reconstruction
Rendering
CD ↓
F-score ↑
NC ↑
PSNR ↑
SSIM ↑
LPIPS ↓
×
×
×
×
19.48
71.21
45.94
19.81
0.881
0.133
✓
×
×
×
6.52
66.34
72.81
20.03
0.860
0.129
✓
✓
×
×
3.91
78.16
78.99
21.32
0.874
0.117
✓
✓
✓
×
4.04
77.48
78.72
20.55
0.866
0.122
✓
✓
✓
✓
3.95
77.87
78.85
24.21
0.907
0.114
Table 3. Ablation study. Geometry (GP) and appearance (AP) priors improve the reconstruction and rendering quality, while adaptive spatial guidance (G-SG and A-SG) further enhances the consistency.
Figure 5 . Comparison of different optimization strategies. Our epsilon-based denoising loss (a), linear noise schedule (b), re-distributed pruning strategy (c), and cosine-power preservation loss weight schedule (d) consistently achieve the most stable convergence.
Method
Runtime ↓
Peak Memory ↓
Avg. Memory ↓
SimRecon
0.18
17.659
13.446
SAM3D
0.13
32.797
22.997
MV-SAM3D
3.13
68.157
51.580
ShapeR
0.37
11.255
5.691
ReconViaGen
0.18
14.911
7.348
CUPID
0.14
13.778
7.123
Table 4. Runtime (hours) and GPU memory (GB) on Replica.
Figure 6 . Comprehensive Reconstruction Results. Our method demonstrates robust, high-quality scene decompositional reconstruction performance on complex real-world captures.
Figure 7 . More Qualitative Comparison. Our method consistently outperforms other methods across different datasets.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Metric
Definition
cd ( cd )
2Accuracy+Completeness
Accuracy
p∈P\mboxmean(p∗∈P∗\mboxmin∣∣p−p∗∣∣1)
Completeness
p∗∈P∗\mboxmean(p∈P\mboxmin∣∣p−p∗∣∣1)
F-score
Precision+Recall2×Precision×Recall
Precision
p∈P\mboxmean(p∗∈P∗\mboxmin∣∣p−p∗∣∣1<0.05)
Recall
p∗∈P∗\mboxmean(p∈P\mboxmin∣∣p−p∗∣∣1<0.05)
Appendix
Table 6 . Evaluation metrics. We list the metrics used to evaluate reconstruction quality and their definitions. P and P∗ are the point clouds sampled from the predicted and ground-truth meshes. np is the normal vector at point p .
Reconstructing the complete geometry of a scene from a single RGB image remains challenging - especially when inferring hidden structures where visual evidence is incomplete. We introduce VolFill, a generative framework that predicts the 3D structure of the complete scene rather than relying on traditional pixel-aligned regression. Our method utilizes a hybrid 3D VAE to compress sparse truncated unsigned distance function grids into a compact latent space, paired with a latent Diffusion Transformer that denoises this representation to recover the complete scene. We condition the generation on geometry foundation models, leveraging rich spatial priors for robust reasoning. Unlike existing methods limited by per-ray constraints or unstructured point-cloud queries, VolFill provides a structured representation that supports direct surface extraction and occupancy queries at scale. Extensive experiments on the SCRREAM and NRGB-D datasets demonstrate that our approach significantly outperforms current baselines, providing a robust foundation for holistic spatial understanding.
In this paper, we introduce \textit{DecoRec}, a novel system designed to elevate single-view 2D images to a decomposed 3D scene mesh. Current methods for single-view scene reconstruction typically rely on object retrieval or the regression of coarse 3D voxels or surfaces, leading to inaccuracies in capturing the appearance and geometry of the input image. The lack of high-quality large-scale scene-level datasets further complicates direct 3D scene generation from single-view images. To achieve high-quality 3D scene generation from a single-view image, DecoRec takes advantage of recent diffusion-based single-view object reconstruction methods to reconstruct individual objects separately. Subsequently, a refinement pipeline is proposed to effectively merge these reconstructed objects, enhancing appearance and geometry through a differentiable rendering technique and diffusion-guided refinement. Our results demonstrate that DecoRec facilitates high-quality single-view scene reconstruction in both geometry and novel synthesis, offering significant benefits for downstream applications like room interior design.
Yuhan Ping, Yuan Liu, Xiaoxiao Long +6
Department of Computer Science, the University of Hong Kong · Department of Computer Science, City University of Hong Kong · Faculty of Humanities and Arts, Macau University of Science and Technology +2
Sparse voxel reconstruction offers an efficient representation for high-fidelity 3D modeling, yet its geometry is commonly optimized from local photometric evidence and discrete visibility statistics. This often leads to fragmented surfaces, excessive subdivision, and floating artifacts, particularly in weakly textured or sparsely observed regions. We introduce SurfSVR, a novel sparse voxel reconstruction paradigm that treats 2D surface priors as explicit 3D geometric regularizers. Instead of directly lifting noisy pixel-wise depth predictions, SurfSVR first organizes each image into coherent surface regions by jointly reasoning over appearance, monocular depth, normals and cross-view geometry. Each region is then represented by an adaptively selected planar or quadratic surface model based on fitting reliability and geometric complexity, while cross-model agreement distinguishes reliable geometry from ambiguous predictions. These structured 2D priors are lifted into 3D and integrated throughout the reconstruction pipeline. They guide surface-adaptive voxel subdivision, provide region-level depth and normal supervision during optimization, enhance geometrically reliable sparse-observed surfaces in voxel pruning, and suppress off-surface floaters during post-refinement training. This unified design converts semantic and geometric coherence in image space into persistent structural constraints in 3D. Extensive experiments on 3 public benchmarks demonstrate that SurfSVR consistently improves sparse voxel reconstruction across scenes with substantially different visibility and geometry characteristics, achieving state-of-the-art reconstruction quality. Codes and models will be released soon.