NeRF and Gaussian splatting methods have been successfully applied on X-ray scenes where the views are too sparse for 3D reconstruction via classical methods. Ultra-sparse scenes with 10 or fewer views such as those with high-rate or low-dose acquisition still, however, present a significant challenge. To address this problem we present a new framework, Optimised X-ray Neural Radiance Fields (OX-NeRF), that combines cross-scene feature learning with scene-specific optimisation to reconstruct sets of related scenes. OX-NeRF employs a convolutional neural network (CNN) to identify cross-scene features while maintaining scene-specific multi-resolution hash grids of spatial features. The paired representations are fused and passed to a multilayer perceptron (MLP); the CNN, hash grids and MLP are then jointly optimised end-to-end. Benchmarking on parallel-beam and cone-beam X-ray datasets shows OX-NeRF provides significantly higher reconstruction accuracy on ultra-sparse scenes compared to existing radiance field methods.
Figures & tables
Figure 1 : Overview of the OX-NeRF pipeline. Left: each scene is recorded by a few X-ray projections at known angles, and a batch of rays is drawn from them by the rule of Section 4.2 . Centre: the scene manager activates the hash grid of the current scene, so the hash encoder reads only that scene’s table, while the shared convolutional encoder reads a pixel-aligned descriptor from the source projections; the fusion module combines the two and a fully fused head returns an attenuation value. Right: attenuation is accumulated along the ray by Eq. 3 and compared with the measurement by mean squared error. Only the hash grids are specific to a scene.
Component
Setting
Hash grid
Per-scene table; 16 levels; 8 features per level; base resolution 16, growth factor 1.5; log2T=12
Image encoder
ResNet-34, first three stages; ImageNet initialisation; P=4 source projections
Rays
2048 rays per iteration; 320 points sampled per ray
Residual sampling
α=1 ; ε=0.25 ; refresh every τ=5 epochs over B=8 scenes
Optimisation
20 000 iterations; Adam, lr 5×10−4 ( 2.5×10−4 for fusion); momentum 0.9, 0.99 [ 25 ]
Table 1 : Hyperparameters. One recipe is used for all datasets. P is the number of source projections; τ and B are the residual-map refresh period and scene count of Section 4.2 ; T is the hash table size.
Figure 2 : Effect of the number of training projections on the Shells dataset. (a) held-out 2D SSIM and (b) 3D SSIM as the budget increases from four to ten projections. (c) OX-NeRF’s prediction of one held-out projection at three budgets on Shells dataset with the measurement for reference.
ONIX [ 51 ]
CombiNeRF [ 6 ]
SAX-NeRF [ 9 ]
R 2 -Gaussian [ 49 ]
OX-NeRF
NVS
3D
NVS
3D
NVS
3D
NVS
3D
NVS
3D
Dataset
Views
SSIM
PSNR
SSIM
SSIM
PSNR
SSIM
SSIM
PSNR
SSIM
SSIM
PSNR
SSIM
SSIM
PSNR
SSIM
Ellipsoids
4
0.7466
17.33
0.3331
0.6494
12.14
0.2434
0.6406
12.36
0.7037
0.8069
16.40
0.6725
0.8893
19.28
0.8406
7
0.7843
20.45
0.3161
0.6247
12.38
0.2508
0.9394
27.31
0.9415
0.8364
19.16
0.7279
0.9915
27.89
0.9566
9
0.8098
20.21
0.2518
0.7360
13.01
0.2555
0.9620
28.84
0.9464
0.9996
22.67
0.7719
0.9966
28.55
0.9613
Shells
4
0.3621
17.12
0.1822
0.3883
18.31
0.3921
0.2916
16.98
0.1816
0.8689
18.87
0.7533
0.8947
20.77
0.7912
Table 2 : Novel view synthesis and 3D reconstruction. NVS reports 2D SSIM on held-out projections; 3D reports PSNR and SSIM of the reconstructed volume against the reference. Best in red bold , second best underlined . 2D PSNR for every cell is in the supplementary material.
Figure 3 : Qualitative 3D reconstruction at ten training views. Each panel is rendered from the reconstructed volume under an identical display window taken from the ground truth, so panels are directly comparable; the inset number is 3D PSNR in dB.
OX-NeRF
R 2 -Gaussian
Views
3D PSNR
3D SSIM
3D PSNR
3D SSIM
15
26.47
0.7442
25.65
0.6984
20
27.04
0.7493
26.97
0.7500
25
27.14
0.7537
27.91
0.7871
Table 4 : High-view Lung CT stress test. The dense Lung CT scene uses the same cone geometry as Table 2 but extends the training budget to 15, 20 and 25 projections. Columns report 3D reconstruction scores as in Table 2 . Best in red bold .
Sparse-view X-ray 3D reconstruction is essential for reducing radiation exposure, but recovering a density field from a handful of X-ray projections is severely ill-posed. Recently, 3D Gaussian Splatting has achieved state-of-the-art performance in sparse-view reconstruction by representing the volume using explicit, optimized primitives, but it requires dozens of projected views. With fewer views, reconstruction quality degrades severely since the explicit primitives are optimized freely without any anatomical information. Anatomical structures, in contrast, share similar geometry and density across a population. Their variations are bounded within a limited range that statistical shape models can capture. This paper proposes a shape-guided Gaussian splatting framework for sparse-view X-ray 3D reconstructions. Our contribution lies in driving Gaussian positions toward anatomically valid configurations, alongside atlas-based density regularization. Our method ensures anatomically consistent reconstruction and improves PSNR by 2.83 dB over a state-of-the-art Gaussian splatting baseline with as few as 5 views. Code Available: https://github.com/polyshape-lab/ShapeGuidedGaussian
Generating high-quality novel views at real-time frame rates remains a central challenge in 3D vision, particularly in sparse-view scenarios. Neural radiance fields have demonstrated robust reconstruction from limited observations, but their reliance on volumetric rendering leads to high computational cost and slow inference. In contrast, Gaussian Splatting methods achieve real-time rendering through rasterization, but their optimization is highly sensitive to the quality of the initial geometry. This sensitivity becomes especially problematic in sparse-view settings, where limited observations often lead to incomplete or noisy point-cloud reconstructions. In this work, we present AugSplat, a simple framework for improving Gaussian Splatting in sparse-view regimes using radiance-field-based view augmentation. We first train a radiance field on the sparse input views and use it to synthesize additional images from nearby novel viewpoints, increasing the effective view-space coverage available for supervision. These synthetic views are then used as auxiliary supervision during Gaussian Splatting optimization. We study two variants: Staged AugSplat, which uses synthetic views for an initial optimization phase before switching to real images, and Dual AugSplat, which jointly trains on real and synthetic views with a decaying synthetic loss weight. Experiments on sparse-view mip-NeRF 360 scenes show that AugSplat improves reconstruction quality over standard Gaussian Splatting. Staged AugSplat achieves the strongest average performance, while Dual AugSplat provides a closely performing formulation that keeps real-image supervision active throughout training, and both variants preserve real-time rendering at inference.
Lorenzo Lazzaroni, Riccardo Bollati, Daniel Barath +2
Remote sensing novel view synthesis under sparse observations remains challenging due to insufficient geometric constraints and limited cross-view supervision. Existing Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) methods are prone to overfitting and face challenges of depth ambiguities, missing cross-view information, and insufficient constraints in under-observed regions. To address these challenges, we propose DIBR-GS, a neural Gaussian Splatting framework that exploits Depth Image-Based Rendering (DIBR) to generate pseudo views for cross-view consistency supervision. Specifically, reliable geometric initialization is constructed by aligning monocular depth priors with sparse SfM reconstruction, and cross-view appearance priors are incorporated into neural Gaussian representations to enhance appearance modeling under sparse observations. Furthermore, we introduce a progressive DIBR-based pseudo-view supervision strategy to provide additional geometric and appearance constraints, enabling more complete reconstruction of weakly observed regions. In addition, a height-constrained anchor growth strategy is designed to suppress unreasonable Gaussian expansion. Experiments demonstrate that the proposed method achieves superior performance over existing approaches when training with only 3 input views. Compared with the previous best-performing method, it improves PSNR by 6.83 dB, with relative gains of 14% in SSIM and 60% in LPIPS, while maintaining competitive computational efficiency. Our code is available at https://github.com/kanehub/DIBR-GS