Organizations: School of Automation and Intelligent Manufacturing, Southern University of Science and Technology · School of Mechatronic Engineering and Automation, Shanghai University · Department of Electrical and Computer Engineering, National University of Singapore · Australian Centre for Robotics, University of Sydney
Partially transmissive screens and protective covers are common in robotic inspection, but they create mixed LiDAR returns from both the foreground material and the scene behind it. Conventional peak-based LiDAR usually discards weak hidden returns, while single-photon LiDAR records time-resolved histograms that preserve attenuated and overlapping echoes. However, existing transient reconstruction methods typically fit a single scene representation to the measured waveform. Under occlusion, weak or nearby foreground--hidden echoes can form a broad peak or subtle shoulder. Because such waveforms can also be explained by a displaced single surface or a thick density distribution, accurate transient fitting does not necessarily imply correct geometry. We propose a state-aware framework for foreground-view and hidden scene reconstruction from occluded single-photon histograms. For each ray, we estimate local echo evidence, identifying no reliable surface evidence, single-return evidence, or two returns. The inferred echo state routes supervision for a two-head neural field: all rays constrain waveform reconstruction, while reliable anchors provide geometry localization. We also introduce a real paired single-photon LiDAR occlusion dataset with occluded and clean captures at fixed poses. Experiments on a real dataset show improved hidden scene depth and point-cloud accuracy over baselines. Our results demonstrate single-photon layered reconstruction as a practical route for 3D perception through partially transmissive occluders.
Figures & tables
Figure 1 : Task overview. An occluder produces mixed foreground and hidden returns along the same ray. Conventional peak-based LiDAR often stops at the occluder, while SPAD transient histograms preserve weak delayed hidden echoes. We use multi-view occluded SPAD measurements to reconstruct separate foreground-view and hidden scene geometry.
Figure 2 : Overview of our state-aware single-photon layered reconstruction framework. From occluded histograms, we infer per-ray echo states and temporal anchors via template fitting, then route waveform, localization, and partition supervision for a two-head neural field. The two heads are rendered independently to recover foreground-view and hidden scene geometry.
Occluder
Method
Foreground
Hidden
Mask-on-Hidden
CD (m) ↓
F1@2 ↑
L1 (m) ↓
CD (m) ↓
F1@2 ↑
L1 (m) ↓
CD (m) ↓
F1@2 ↑
L1 (m) ↓
Black mesh
DyNFL
0.373
0.692
1.656
0.408
0.642
1.904
0.375
0.499
1.197
Transientangelo
0.194
0.841
1.768
0.155
0.857
1.244
0.495
0.525
0.840
TransientNeRF
0.299
0.748
0.986
0.179
0.838
1.092
0.415
0.546
0.967
Flying with Photons
0.252
0.634
1.064
0.241
0.705
1.056
0.578
0.453
1.386
Multilayer Imaging
0.076
0.952
0.427
0.134
0.846
0.873
0.245
0.491
0.642
Table 1 : Quantitative comparison on the captured multi-view occlusion LiDAR dataset. CD and L1 are lower better, while F1@2 is higher better. Best results within each occluder group are in bold, and second-best results are underlined.
Figure 3 : Qualitative comparison of foreground-view and hidden scene reconstructions across methods.
Components
Foreground
Hidden
Mask-on-Hidden
State
Lloc
Lsep
F1@2 ↑
L1 (m) ↓
F1@2 ↑
L1 (m) ↓
F1@2 ↑
L1 (m) ↓
✓
0.873
0.430
0.568
2.258
0.291
1.373
✓
✓
0.852
0.557
0.926
0.427
0.919
0.419
✓
✓
0.884
0.402
0.904
0.493
0.864
0.428
✓
✓
✓
0.901
0.419
0.916
0.459
0.863
0.415
Table 2: Ablation study. Best results are in bold.
Figure 4 : Sensitivity to loss weights. Foreground and hidden Depth L1 errors are reported when varying λloc and λsep .
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Occluder
Method
Foreground
Hidden
Mask-on-Hidden
CD (m) ↓
F1@2 ↑
L1 (m) ↓
CD (m) ↓
F1@2 ↑
L1 (m) ↓
CD (m) ↓
F1@2 ↑
L1 (m) ↓
Black mesh
DyNFL
0.373 ± 0.258
0.692 ± 0.152
1.656 ± 0.869
0.408 ± 0.236
0.642 ± 0.130
1.904 ± 0.762
0.375 ± 0.179
0.499 ± 0.079
1.197 ± 0.594
Transientangelo
0.194 ± 0.057
0.841 ± 0.039
1.768 ± 0.238
0.155 ± 0.025
0.857 ± 0.013
1.244 ± 0.179
0.495 ± 0.187
0.525 ± 0.082
0.840 ± 0.259
TransientNeRF
0.299 ± 0.120
0.748 ± 0.085
0.986 ± 0.334
0.179 ± 0.050
0.838 ± 0.034
1.092 ± 0.242
0.415 ± 0.198
0.546 ± 0.112
0.967 ± 0.306
Flying with Photons
0.252 ± 0.102
0.634 ± 0.086
1.064 ± 0.259
0.241 ± 0.061
0.705 ± 0.053
1.056 ± 0.184
0.578 ± 0.183
0.453 ± 0.075
1.386 ± 0.122
Multilayer Imaging
0.076 ± 0.001
0.952 ± 0.001
0.427 ± 0.007
0.134 ± 0.006
0.846 ± 0.008
0.873 ± 0.029
0.245 ± 0.037
0.491 ± 0.105
0.642 ± 0.096
Appendix
Table 3 : Expanded quantitative results with additional baselines and standard deviations. Results are grouped by occluder type and reported as mean ± std over scenes in each group. The best and second-best results within each occluder block are bolded and underlined, respectively. MLI denotes Multilayer Imaging, and FwP denotes Flying with Photons.
First Photon
Foreground
Hidden
Mask-on-Hidden
F1@2 ↑
L1 (m) ↓
CD (m) ↓
F1@2 ↑
L1 (m) ↓
CD (m) ↓
F1@2 ↑
L1 (m) ↓
CD (m) ↓
w/o
0.9449±0.0101
0.3685±0.0450
0.0761±0.0137
0.8962±0.0346
0.4683±0.0750
0.1165±0.0195
0.8629±0.0600
0.4495±0.2960
0.1407±0.0506
w
0.901±0.057
0.419±0.080
0.104±0.032
0.917±0.010
0.459±0.040
0.102±0.005
0.861±0.069
0.421±0.193
0.142±0.050
Appendix
Table 4: Ablation Study of the First Photon Model.
Material
State Acc.
Single Recall
Two Recall
FG MAE
Hidden MAE
Blacknet
93.9%
80.0%
96.4%
0.63
0.93
Mosquito net
93.9%
80.0%
96.4%
0.93
2.44
Shower
100.0%
100.0%
100.0%
4.24
5.66
Overall
96.0%
86.7%
97.6%
1.99
3.07
Appendix
Table 5 : Accuracy of echo-state classification and foreground/background anchor localization across occluding materials. Anchor localization errors are reported in temporal bins.
Foreground
Hidden
Mask-on-Hidden
F1@2 ↑
L1 (m) ↓
CD (m) ↓
F1@2 ↑
L1 (m) ↓
CD (m) ↓
F1@2 ↑
L1 (m) ↓
CD (m) ↓
Anchor Projection
0.765±0.020
1.010±0.100
0.195±0.020
0.689±0.012
1.075±0.053
0.238±0.010
0.353±0.067
0.947±0.110
0.442±0.053
Ours
0.901±0.057
0.419±0.080
0.104±0.032
0.917±0.010
0.459±0.040
0.102±0.005
0.861±0.069
0.421±0.193
0.142±0.050
Appendix
Table 6: Cross-view projection accuracy of extracted temporal anchors. We directly back-project training-view anchors into foreground-view and hidden-scene point clouds, project them to test views, and evaluate with the reconstruction metrics.
Figure 5 : Automatically estimated foreground/visible response templates κfg from first echoes of different occluding materials. The profiles are temporally aligned and peak-normalized for shape comparison.
Figure 6 : Automatically estimated hidden-response templates κhid from second echoes of different occluding materials. Compared with the first-echo templates, the hidden responses are broader and have longer temporal tails.
Figure 7 : Material- and location-dependent transient waveforms under occlusion. For each occluding material, we mark four representative spatial locations and show their measured transient histograms. Different locations produce different echo structures, ranging from narrow single returns to broadened, asymmetric, or clearly separated multi-return waveforms. The waveform shape also changes with the occluding material.
Figure 8 : Additional qualitative comparison on the duck scene under different foreground occluders. Depth maps and depth-colored point clouds are visualized within 0–15 m, while mask-on-hidden depth-error maps are visualized within 0–2 m.
Figure 9 : Additional qualitative comparison on the house scene under different foreground occluders. Depth maps and depth-colored point clouds are visualized within 0–15 m, while mask-on-hidden depth-error maps are visualized within 0–2 m.
Figure 10 : Additional qualitative comparison on the triangle scene under different foreground occluders. Depth maps and depth-colored point clouds are visualized within 0–15 m, while mask-on-hidden depth-error maps are visualized within 0–2 m.