NeRF and Gaussian splatting methods have been successfully applied on X-ray scenes where the views are too sparse for 3D reconstruction via classical methods. Ultra-sparse scenes with 10 or fewer views such as those with high-rate or low-dose acquisition still, however, present a significant challenge. To address this problem we present a new framework, Optimised X-ray Neural Radiance Fields (OX-NeRF), that combines cross-scene feature learning with scene-specific optimisation to reconstruct sets of related scenes. OX-NeRF employs a convolutional neural network (CNN) to identify cross-scene features while maintaining scene-specific multi-resolution hash grids of spatial features. The paired representations are fused and passed to a multilayer perceptron (MLP); the CNN, hash grids and MLP are then jointly optimised end-to-end. Benchmarking on parallel-beam and cone-beam X-ray datasets shows OX-NeRF provides significantly higher reconstruction accuracy on ultra-sparse scenes compared to existing radiance field methods.
Figures & tables
Figure 1 : Overview of the OX-NeRF pipeline. Left: each scene is recorded by a few X-ray projections at known angles, and a batch of rays is drawn from them by the rule of Section 4.2 . Centre: the scene manager activates the hash grid of the current scene, so the hash encoder reads only that scene’s table, while the shared convolutional encoder reads a pixel-aligned descriptor from the source projections; the fusion module combines the two and a fully fused head returns an attenuation value. Right: attenuation is accumulated along the ray by Eq. 3 and compared with the measurement by mean squared error. Only the hash grids are specific to a scene.
Component
Setting
Hash grid
Per-scene table; 16 levels; 8 features per level; base resolution 16, growth factor 1.5; log2T=12
Image encoder
ResNet-34, first three stages; ImageNet initialisation; P=4 source projections
Rays
2048 rays per iteration; 320 points sampled per ray
Residual sampling
α=1 ; ε=0.25 ; refresh every τ=5 epochs over B=8 scenes
Optimisation
20 000 iterations; Adam, lr 5×10−4 ( 2.5×10−4 for fusion); momentum 0.9, 0.99 [ 25 ]
Table 1 : Hyperparameters. One recipe is used for all datasets. P is the number of source projections; τ and B are the residual-map refresh period and scene count of Section 4.2 ; T is the hash table size.
Figure 2 : Effect of the number of training projections on the Shells dataset. (a) held-out 2D SSIM and (b) 3D SSIM as the budget increases from four to ten projections. (c) OX-NeRF’s prediction of one held-out projection at three budgets on Shells dataset with the measurement for reference.
ONIX [ 51 ]
CombiNeRF [ 6 ]
SAX-NeRF [ 9 ]
R 2 -Gaussian [ 49 ]
OX-NeRF
NVS
3D
NVS
3D
NVS
3D
NVS
3D
NVS
3D
Dataset
Views
SSIM
PSNR
SSIM
SSIM
PSNR
SSIM
SSIM
PSNR
SSIM
SSIM
PSNR
SSIM
SSIM
PSNR
SSIM
Ellipsoids
4
0.7466
17.33
0.3331
0.6494
12.14
0.2434
0.6406
12.36
0.7037
0.8069
16.40
0.6725
0.8893
19.28
0.8406
7
0.7843
20.45
0.3161
0.6247
12.38
0.2508
0.9394
27.31
0.9415
0.8364
19.16
0.7279
0.9915
27.89
0.9566
9
0.8098
20.21
0.2518
0.7360
13.01
0.2555
0.9620
28.84
0.9464
0.9996
22.67
0.7719
0.9966
28.55
0.9613
Shells
4
0.3621
17.12
0.1822
0.3883
18.31
0.3921
0.2916
16.98
0.1816
0.8689
18.87
0.7533
0.8947
20.77
0.7912
Table 2 : Novel view synthesis and 3D reconstruction. NVS reports 2D SSIM on held-out projections; 3D reports PSNR and SSIM of the reconstructed volume against the reference. Best in red bold , second best underlined . 2D PSNR for every cell is in the supplementary material.
Figure 3 : Qualitative 3D reconstruction at ten training views. Each panel is rendered from the reconstructed volume under an identical display window taken from the ground truth, so panels are directly comparable; the inset number is 3D PSNR in dB.
OX-NeRF
R 2 -Gaussian
Views
3D PSNR
3D SSIM
3D PSNR
3D SSIM
15
26.47
0.7442
25.65
0.6984
20
27.04
0.7493
26.97
0.7500
25
27.14
0.7537
27.91
0.7871
Table 4 : High-view Lung CT stress test. The dense Lung CT scene uses the same cone geometry as Table 2 but extends the training budget to 15, 20 and 25 projections. Columns report 3D reconstruction scores as in Table 2 . Best in red bold .