Shadow Reduction in Ultrasound Imaging Using Differentiable Simulation and Radiance Field Decomposition
Authors: Valentin Bacher, Pak Hei Yeung, Bernhard Kainz, Madeleine K. Wyburd, Nicola K. Dinsdale, Michael Gray, Ana I. L. Namburete
Organizations: Oxford Machine Learning in NeuroImaging Lab, University of Oxford, United Kingdom · Oxford Machine Learning in NeuroImaging Lab, Department of Computer Science, University of Oxford, OX1 3QD, United Kingdom · Quantitative Healthcare Analysis · Quantitative Healthcare Analysis (qurAI) Group, Informatics Institute, University of Amsterdam, Amsterdam, 1098 XH, The Netherlands · Friedrich-Alexander-Universität Erlangen-Nürnberg, Germany · Imperial College London, United Kingdom · Imperial College London, 180 Queen’s Gate, London, SW7 2AZ, United Kingdom · Department of Computer Science, University of Copenhagen, Denmark · Institute of Biomedical Engineering, University of Oxford, United Kingdom · Institute of Biomedical Engineering, University of Oxford, Marcela Botnar Wing, Oxford, OX3 7LD, United Kingdom
Acoustic shadows from bone and other highly attenuating tissues obscure clinically important structures in ultrasound. In fetal brain imaging, skull-induced artefacts disproportionately degrade the hemisphere closer to the transducer (proximal), limiting symmetric assessment of the two hemispheres. Existing correction methods require raw scanner data, impose restrictive assumptions on tissue properties, or rely on generative models that may hallucinate anatomy. We present RFlash, a physics-informed post-processing method that decomposes beamformed ultrasound images into explicit attenuation and scatter-intensity maps using a differentiable radiance-field formulation of image formation. Attenuation-adaptive re-rendering then removes the dependence of the signal at each depth on the intervening tissue, equivalent to virtually advancing the transducer into the tissue. Across 1,261 3D fetal brain volumes, 143 real 2D curvilinear abdominal scans, and 1,200 simulated 2D linear-probe liver scans, RFlash reduces shadow-related intensity differences more effectively than classical Hughes-Duck attenuation correction. For a gestational-age model trained on the distal hemisphere (further from the transducer) and applied to the proximal hemisphere, prediction error decreases by 5.1 days (40%) relative to the original images. The estimated attenuation maps also yield shadow-confidence maps that improve random-forest bone-shadow segmentation over the image alone and receive greater SHAP importance than an existing neural confidence-map baseline, suggesting greater physical consistency. RFlash requires neither hardware modification nor access to raw scanner data and supports 2D and 3D acquisitions with linear and curvilinear probes, making it widely applicable allowing clinicians to use our method on their already acquired scanners and images.
Figures & tables
Figure 1 : Acoustic artefacts in clinical B-mode ultrasound. Solid red boxes mark acoustic shadows behind strongly attenuating structures. The dashed blue box marks reverberation, the dashed green box marks enhancement artefacts, and the dashed orange box marks the contrast difference between the proximal (closer to the transducer) and distal (further away from the transducer) fetal brain hemispheres. RFlash targets enhancement and shadow artefacts and, to some extent, the contrast difference between the hemispheres.
Figure 2 : Principle of shadow-free rendering. Left: in a conventional acquisition, the intensity at a given depth (green ray) depends on all tissue between the transducer and that depth, so attenuating structures such as the skull cast shadows into everything below them. Right: RFlash renders each depth as though the transducer had been advanced to just above it (orange arcs, short green rays), making the rendered intensity independent of overlying tissue and suppressing the shadow.
Figure 3 : The two phases of RFlash . (a) Decomposition: attenuation A and scatter intensity Φ are stored in an explicit volume, sampled along scanlines by trilinear interpolation, and passed through the differentiable forward simulator K:A,Φ↦Y^ . The rendered slice Y^ is compared with the original slice Y by the loss l(Y,Y^) , and gradients (dashed arrows) update the parameter volumes. (b) Shadow-reduced rendering (bottom): the converged Φ from (a) is re-rendered through the shadow-avoiding operator K′:Φ↦Y′ , which omits the attenuation term and so produces a shadow-free image. Solid arrows: forward simulation. Dashed arrows: gradient flow.
Figure 4 : The canonical ultrasound imaging pipeline on which RFlash is built. Tissue response (left) is modulated by the propagating sound field, then processed in the scanner (centre: envelope detection, time-gain compensation, log compression) and converted to a display image (right: A/D conversion, binning, scan conversion). RFlash operates on the displayed, beamformed image and inverts the shaded operations. A/D conversion and binning are lossy and are not modelled. Compare with Figure 6 , which shows the corresponding differentiable pipeline for beamformed images.
Figure 5 : Geometry of the 3D acquisition. A mechanically swept probe is modelled as a curvilinear 2D array (yellow fan) rotated about an internal pivot axis (dashed), so that a voxel is addressed by depth r and two angular coordinates: φ within the fan and ϑ about the pivot. Each scanline is treated as an independent propagation path.
Figure 6 : Visual representation of the differentiable simulation pipeline as defined by eq. 1 – eq. 10 . The first target (top right) is the simulated ultrasound image Y^ , while the lower two targets are the shadow-reduced image Y′ and the shadow map M , which are only relevant after the tissue parameters have been estimated (b in Figure 3 ).
Figure 7 : Qualitative comparison across acquisition geometries: fetal brain (3D mechanically swept), abdominal (curvilinear 2D) and simulated liver (linear 2D). Rows: original image, Hughes shadow reduction, RFlash . Orange arrows mark regions of interest: in the fetal brain, recovery of the Sylvian fissure in the proximal hemisphere. In the abdominal, over-brightening of shadowed regions by Hughes reduction, producing cloudy artefacts that RFlash avoids. In the liver dataset, enhancement artefacts, Hughes method does not remove.
Dataset
ΔμO↓
ΔμH↓
ΔμSR↓
(ΔμO−ΔμSR)
(ΔμH−ΔμSR)
Syn liver
64.12±27.21
13.61±7.37
16.93±17.47
47.19±20.06∗
−3.32±18.40∗
Abdomen
87.65±16.84
13.41±10.13
8.04±6.41
79.61±18.41∗
5.37±10.33∗
Fetal Brain
129.04±7.59
48.26±22.84
41.94±11.73
87.10±12.57∗
6.32±29.55∗
Table 1 : Mean intensity difference Δμ across datasets for the original ( O ), Hughes shadow-reduced ( H ), and RFlash ( SR ) volumes. Values are shown as mean ± standard deviation. The two right-most columns report the subject-wise differences relative to RFlash . Differences marked with ∗ are significantly different from 0 in a t-test with a p-value <60.01=0.001667 for multiple-comparison correction.
Hemisphere
ϵo
ϵH
ϵSR
ϵO−ϵSR
ϵH−ϵSR
Distal
4.70±3.96
4.27±3.24
3.96±2.89
0.75[0.27,1.22]
0.31[−0.00,0.64]
Proximal
12.66±8.76
9.44±7.81
7.60±5.73
5.08[4.00,6.15]∗
1.87[0.76,2.96]∗
Table 2 : Average prediction error ϵ in days for fetal age prediction. Training and validation were performed on the distal hemispheres, while testing was performed on the distal and proximal hemispheres. Prediction errors ϵO , ϵH , and ϵSR are reported as mean ± standard deviation. Paired differences are reported as mean [95% fetus-level cluster-bootstrap confidence interval], accounting for fetuses scanned at multiple gestational ages. Statistics were calculated over 263 scans from 252 fetuses. Statistical significance was assessed using a two-sided Wilcoxon signed-rank test, with differences from multiple scans of the same fetus first averaged so that each fetus contributed one independent observation. Asterisks mark mean paired differences significantly different from zero according to the Wilcoxon test after Bonferroni correction ( p<40.01=0.0025 ).
Figure 8 : Bland–Altman plots for age prediction on both hemispheres across the three datasets: the original images (orig), the images with Hughes shadow reduction (H), and the images with RFlash shadow reduction (SR). For readability, the first two datasets show only the convex hull, mean, and standard deviation. The full results are provided in Section A.2 .
Figure 9 : Representative cases from the age-prediction experiment. Top row: a case where shadow reduction substantially increased contrast in the proximal hemisphere. The green arrow marks the cerebellum, which is visibly better resolved after shadow reduction. Bottom row: a case where a residual grey-value gradient across the hemispheres persists (red box) in the original and Hughes-reduced images but is largely removed by RFlash .
Figure 10 : Overlay comparison of the shadow confidence maps for two examples. The columns compare the overlays obtained from the method of Qingjie Meng et al. [ 57 ] and our method.
Name
Accuracy ↑
Precision ↑
Dice ↑
Hausdorff dist. ↓
Baseline
0.912±0.006
0.585±0.001
0.532±0.003
8.196±0.034
Yesilkaynak
0.922±0.006
0.699±0.002
0.519±0.005
5.404±0.188
Ours
0.930±0.005
0.693±0.001
0.588±0.004
5.104±0.107
Table 3: Performance summary reported to three decimal places, averaged over 100 random forests. The baseline row used the ultrasound image as input, the second row used the shadow maps from Yesilkaynak [ 58 ] , and the third row used the shadow maps from our method. The best value in each column is highlighted in bold.
Figure 11 : Visual comparison for three selected ultrasound simulations of a liver phantom. The first column shows the ultrasound images with corresponding bone-shadow maps: the first row uses the ground-truth map, while the other rows use random-forest predictions. Yesil. refers to the methods introduced by [ 58 ] . The last two columns show confidence maps with the ground-truth outlines.
Figure 12 : Explainability maps for a single random forest computed from 1000 randomly selected paths using TreeExplainer. The columns correspond to the baseline model using the ultrasound image only, the model using the Yesilkaynak confidence map [ 58 ] , and our model. The first row shows SHAP feature-importance maps for the ultrasound images, the middle row shows the corresponding colour bars, and the final row shows the feature maps when available. The feature maps show the average importance of the 5×5 -pixel patch around the classified centre pixel. The colours are matched within each random tree, that is, within each column. Higher SHAP magnitude on the RFlash confidence map (right) indicates that the random forest relied on it more than on the baseline map.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
ΔμO
ΔμH
ΔμSR
(ΔμO−ΔμSR)
(ΔμH−ΔμSR)
Syn liver
64.12±27.21
13.44±24.46
16.93±17.47
47.19±20.06∗
−3.49±8.75∗
Abdomen
87.65±16.84
27.91±11.24
8.04±6.41
79.61±18.41∗
19.87±11.25∗
Fetal Brain
129.04±7.59
58.90±18.16
41.94±11.73
87.10±12.57∗
16.95±19.33∗
Appendix
Table 4 : Mean intensity difference Δμ across datasets for the original ( O ), Hughes shadow-reduced ( H ), and RFlash ( SR ) volumes when all measures are computed with the RFlash shadow maps. Values are shown as mean ± standard deviation. The two right-most columns report the subject-wise differences relative to RFlash . Differences marked with ∗ are significantly different from 0 in a t-test with a p-value <60.01=0.001667 for multiple-comparison correction.
Dataset
ΔμO
ΔμH
ΔμSR
(ΔμO−ΔμSR)
(ΔμH−ΔμSR)
Syn liver
62.23±9.37
13.61±7.37
22.21±5.80
40.02±6.35∗
−8.59±8.81∗
Abdomen
70.26±22.04
13.41±10.13
11.48±7.90
58.78±21.19∗
1.93±10.99
Fetal Brain
88.01±12.29
48.26±22.84
15.25±11.92
72.75±15.33∗
33.00±24.52∗
Appendix
Table 5 : Mean intensity difference Δμ across datasets for the original ( O ), Hughes shadow-reduced ( H ), and RFlash ( SR ) volumes when all measures are computed with the Hughes shadow maps. Values are shown as mean ± standard deviation. The two right-most columns report the subject-wise differences relative to RFlash . Differences marked with ∗ are significantly different from 0 in a t-test with a p-value <60.01=0.001667 for multiple-comparison correction.
Figure 13 : This figure complements Figure 8 by showing the individual data points underlying the convex hulls in that figure.
Figure 14 : Similarity measures between simulated and original volumes as a function of n , where n is the edge length of the sampled cube in simulation space. The red line indicates the chosen value.
Figure 15 : This figure shows example volumes reconstructed using different densities of sampling in simulation space.
Figure 16 : Heat maps of shadow-reduction measures for different combinations of initial values. The red box marks the selected combination.
Figure 17 : SSIM and L2 after convergence for different weighting factors. The red line indicates the chosen value.
Figure 18 : Convergence of the L2 and SSIM metrics over epochs. The red line indicates the chosen value.
Figure 19 : Mean and confidence intervals for the three metrics described above as a function of the number of buckets used to calculate the histograms. The red line indicates the chosen value.