Organizations: School of Automation and Intelligent Manufacturing, Southern University of Science and Technology · School of Mechatronic Engineering and Automation, Shanghai University · Department of Electrical and Computer Engineering, National University of Singapore · Australian Centre for Robotics, University of Sydney
Partially transmissive screens and protective covers are common in robotic inspection, but they create mixed LiDAR returns from both the foreground material and the scene behind it. Conventional peak-based LiDAR usually discards weak hidden returns, while single-photon LiDAR records time-resolved histograms that preserve attenuated and overlapping echoes. However, existing transient reconstruction methods typically fit a single scene representation to the measured waveform. Under occlusion, weak or nearby foreground--hidden echoes can form a broad peak or subtle shoulder. Because such waveforms can also be explained by a displaced single surface or a thick density distribution, accurate transient fitting does not necessarily imply correct geometry. We propose a state-aware framework for foreground-view and hidden scene reconstruction from occluded single-photon histograms. For each ray, we estimate local echo evidence, identifying no reliable surface evidence, single-return evidence, or two returns. The inferred echo state routes supervision for a two-head neural field: all rays constrain waveform reconstruction, while reliable anchors provide geometry localization. We also introduce a real paired single-photon LiDAR occlusion dataset with occluded and clean captures at fixed poses. Experiments on a real dataset show improved hidden scene depth and point-cloud accuracy over baselines. Our results demonstrate single-photon layered reconstruction as a practical route for 3D perception through partially transmissive occluders.
Figures & tables
Figure 1 : Task overview. An occluder produces mixed foreground and hidden returns along the same ray. Conventional peak-based LiDAR often stops at the occluder, while SPAD transient histograms preserve weak delayed hidden echoes. We use multi-view occluded SPAD measurements to reconstruct separate foreground-view and hidden scene geometry.
Figure 2 : Overview of our state-aware single-photon layered reconstruction framework. From occluded histograms, we infer per-ray echo states and temporal anchors via template fitting, then route waveform, localization, and partition supervision for a two-head neural field. The two heads are rendered independently to recover foreground-view and hidden scene geometry.
Occluder
Method
Foreground
Hidden
Mask-on-Hidden
CD (m) ↓
F1@2 ↑
L1 (m) ↓
CD (m) ↓
F1@2 ↑
L1 (m) ↓
CD (m) ↓
F1@2 ↑
L1 (m) ↓
Black mesh
DyNFL
0.373
0.692
1.656
0.408
0.642
1.904
0.375
0.499
1.197
Transientangelo
0.194
0.841
1.768
0.155
0.857
1.244
0.495
0.525
0.840
TransientNeRF
0.299
0.748
0.986
0.179
0.838
1.092
0.415
0.546
0.967
Flying with Photons
0.252
0.634
1.064
0.241
0.705
1.056
0.578
0.453
1.386
Multilayer Imaging
0.076
0.952
0.427
0.134
0.846
0.873
0.245
0.491
0.642
Table 1 : Quantitative comparison on the captured multi-view occlusion LiDAR dataset. CD and L1 are lower better, while F1@2 is higher better. Best results within each occluder group are in bold, and second-best results are underlined.
Figure 3 : Qualitative comparison of foreground-view and hidden scene reconstructions across methods.
Components
Foreground
Hidden
Mask-on-Hidden
State
Lloc
Lsep
F1@2 ↑
L1 (m) ↓
F1@2 ↑
L1 (m) ↓
F1@2 ↑
L1 (m) ↓
✓
0.873
0.430
0.568
2.258
0.291
1.373
✓
✓
0.852
0.557
0.926
0.427
0.919
0.419
✓
✓
0.884
0.402
0.904
0.493
0.864
0.428
✓
✓
✓
0.901
0.419
0.916
0.459
0.863
0.415
Table 2: Ablation study. Best results are in bold.
Figure 4 : Sensitivity to loss weights. Foreground and hidden Depth L1 errors are reported when varying λloc and λsep .
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Occluder
Method
Foreground
Hidden
Mask-on-Hidden
CD (m) ↓
F1@2 ↑
L1 (m) ↓
CD (m) ↓
F1@2 ↑
L1 (m) ↓
CD (m) ↓
F1@2 ↑
L1 (m) ↓
Black mesh
DyNFL
0.373 ± 0.258
0.692 ± 0.152
1.656 ± 0.869
0.408 ± 0.236
0.642 ± 0.130
1.904 ± 0.762
0.375 ± 0.179
0.499 ± 0.079
1.197 ± 0.594
Transientangelo
0.194 ± 0.057
0.841 ± 0.039
1.768 ± 0.238
0.155 ± 0.025
0.857 ± 0.013
1.244 ± 0.179
0.495 ± 0.187
0.525 ± 0.082
0.840 ± 0.259
TransientNeRF
0.299 ± 0.120
0.748 ± 0.085
0.986 ± 0.334
0.179 ± 0.050
0.838 ± 0.034
1.092 ± 0.242
0.415 ± 0.198
0.546 ± 0.112
0.967 ± 0.306
Flying with Photons
0.252 ± 0.102
0.634 ± 0.086
1.064 ± 0.259
0.241 ± 0.061
0.705 ± 0.053
1.056 ± 0.184
0.578 ± 0.183
0.453 ± 0.075
1.386 ± 0.122
Multilayer Imaging
0.076 ± 0.001
0.952 ± 0.001
0.427 ± 0.007
0.134 ± 0.006
0.846 ± 0.008
0.873 ± 0.029
0.245 ± 0.037
0.491 ± 0.105
0.642 ± 0.096
Appendix
Table 3 : Expanded quantitative results with additional baselines and standard deviations. Results are grouped by occluder type and reported as mean ± std over scenes in each group. The best and second-best results within each occluder block are bolded and underlined, respectively. MLI denotes Multilayer Imaging, and FwP denotes Flying with Photons.
First Photon
Foreground
Hidden
Mask-on-Hidden
F1@2 ↑
L1 (m) ↓
CD (m) ↓
F1@2 ↑
L1 (m) ↓
CD (m) ↓
F1@2 ↑
L1 (m) ↓
CD (m) ↓
w/o
0.9449±0.0101
0.3685±0.0450
0.0761±0.0137
0.8962±0.0346
0.4683±0.0750
0.1165±0.0195
0.8629±0.0600
0.4495±0.2960
0.1407±0.0506
w
0.901±0.057
0.419±0.080
0.104±0.032
0.917±0.010
0.459±0.040
0.102±0.005
0.861±0.069
0.421±0.193
0.142±0.050
Appendix
Table 4: Ablation Study of the First Photon Model.
Material
State Acc.
Single Recall
Two Recall
FG MAE
Hidden MAE
Blacknet
93.9%
80.0%
96.4%
0.63
0.93
Mosquito net
93.9%
80.0%
96.4%
0.93
2.44
Shower
100.0%
100.0%
100.0%
4.24
5.66
Overall
96.0%
86.7%
97.6%
1.99
3.07
Appendix
Table 5 : Accuracy of echo-state classification and foreground/background anchor localization across occluding materials. Anchor localization errors are reported in temporal bins.
Foreground
Hidden
Mask-on-Hidden
F1@2 ↑
L1 (m) ↓
CD (m) ↓
F1@2 ↑
L1 (m) ↓
CD (m) ↓
F1@2 ↑
L1 (m) ↓
CD (m) ↓
Anchor Projection
0.765±0.020
1.010±0.100
0.195±0.020
0.689±0.012
1.075±0.053
0.238±0.010
0.353±0.067
0.947±0.110
0.442±0.053
Ours
0.901±0.057
0.419±0.080
0.104±0.032
0.917±0.010
0.459±0.040
0.102±0.005
0.861±0.069
0.421±0.193
0.142±0.050
Appendix
Table 6: Cross-view projection accuracy of extracted temporal anchors. We directly back-project training-view anchors into foreground-view and hidden-scene point clouds, project them to test views, and evaluate with the reconstruction metrics.
Figure 5 : Automatically estimated foreground/visible response templates κfg from first echoes of different occluding materials. The profiles are temporally aligned and peak-normalized for shape comparison.
Figure 6 : Automatically estimated hidden-response templates κhid from second echoes of different occluding materials. Compared with the first-echo templates, the hidden responses are broader and have longer temporal tails.
Figure 7 : Material- and location-dependent transient waveforms under occlusion. For each occluding material, we mark four representative spatial locations and show their measured transient histograms. Different locations produce different echo structures, ranging from narrow single returns to broadened, asymmetric, or clearly separated multi-return waveforms. The waveform shape also changes with the occluding material.
Figure 8 : Additional qualitative comparison on the duck scene under different foreground occluders. Depth maps and depth-colored point clouds are visualized within 0–15 m, while mask-on-hidden depth-error maps are visualized within 0–2 m.
Figure 9 : Additional qualitative comparison on the house scene under different foreground occluders. Depth maps and depth-colored point clouds are visualized within 0–15 m, while mask-on-hidden depth-error maps are visualized within 0–2 m.
Figure 10 : Additional qualitative comparison on the triangle scene under different foreground occluders. Depth maps and depth-colored point clouds are visualized within 0–15 m, while mask-on-hidden depth-error maps are visualized within 0–2 m.
Single-photon LiDAR (SPL) based on single-photon avalanche diode (SPAD) sensing enables time-resolved photon measurements with extreme sensitivity, offering unique potential for active 3D perception in photon-starved scenarios.However, real-world single photon perception remains fundamentally challenging due to unique measurement noise and complex multi-return transient phenomena, which jointly complicate geometric reconstruction and semantic scene understanding. Despite growing interest in SPAD-based sensing, existing studies are largely limited to simulated data or small-scale controlled captures. As a result, systematic evaluation of real-world single photon perception across depth estimation, multi-view reconstruction, and 3D semantic understanding remains underexplored. To bridge this gap, we introduce SP-TransientBench (STB), a real-captured multi-task benchmark for single photon perception. SP-TransientBenc comprises 10 diverse scenes and 10,297 views captured using a solid-state single-photon LiDAR at 256×192 resolution. Each view provides full time-of-flight histograms with multi-return behavior,standardized metadata, and calibrated camera poses for multi-view evaluation. We further provide 13-class 3D semantic annotations for selected scenes. By providing dedicated data splits and evaluation protocols for each task, STB enables consistent and reproducible benchmarking of real-world single photon perception across multiple 3D vision problems. The dataset and code will be released upon acceptance.
Hongzhou Dong, Zili Zhang, Ziting Wen +12
Shanghai University, Shanghai, China · Southern University of Science and Technology, Shenzhen, China · The University of Sydney, Sydney, Australia
Occlusion-robust scene recovery remains a major challenge in computational imaging, particularly in natural environments where dense foreground vegetation severely limits visibility. We propose a vision-reasoning-guided light field occlusion removal framework that combines the visibility recovery capability of light field integration (LFI) with the semantic reasoning capacity of vision-language models (VLMs). Multi-view observations are first integrated via LFI to suppress foreground occlusions and produce an initial visibility-enhanced representation. A VLM is then incorporated as a conditional semantic prior to restore degraded structures and recover fine details, guided by the observed measurements. To improve recovery consistency and reduce hallucination artifacts, we introduce a multi-sample fusion strategy that aggregates multiple generated hypotheses into a unified estimate. Experimental results on synthetic and real-world datasets demonstrate state-of-the-art performance, achieving the highest average SSIM across four synthetic light field benchmark scenes (4-Syn) and strong generalization across structured and unstructured acquisition settings. These results highlight the effectiveness of combining physical imaging constraints with vision-language reasoning for robust perception under severe occlusion, with applicability to search-and-rescue and exploratory robotic navigation.
Mohamed Youssef, Oliver Bimber
Department of Computer Science, Johannes Kepler University, Altenbergerstr. 65, Linz, 4040, Austria
Consumer LiDARs in mobile devices and robots typically output a single depth value per pixel. Yet internally, they record full time-resolved histograms containing direct and multi-bounce light returns; these multi-bounce returns encode rich non-line-of-sight (NLOS) cues that can enable perception of hidden objects in a scene. However, severe hardware limitations of consumer LiDARs make NLOS reconstruction with conventional methods difficult. In this work, we motivate a complementary direction: enabling NLOS perception with low-cost LiDARs through data-driven inference. We present DENALI, the first large-scale real-world dataset of space-time histograms from low-cost LiDARs capturing hidden objects. We capture time-resolved LiDAR histograms for 72,000 hidden-object scenes across diverse object shapes, positions, lighting conditions, and spatial resolutions. Using our dataset, we show that consumer LiDARs can enable accurate, data-driven NLOS perception. We further identify key scene and modeling factors that limit performance, as well as simulation-fidelity gaps that hinder current sim-to-real transfer, motivating future work toward scalable NLOS vision with consumer LiDARs.
Nikhil Behari, Diego Rivero, Luke Apostolides +3
Massachusetts Institute of Technology · Technische Universität Berlin