Latent LiDAR pipelines suffer from flying pixels: convolutional VAEs blur sharp radial depth discontinuities, yielding edge depths that back-project to points floating between surfaces. We identify this as a major, directly correctable decoder bottleneck and introduce CRISP: a pixel-space diffusion decoder with a backbone-agnostic latent adapter, DiT-based denoiser, and support mask predictor. CRISP replaces video-VAE and LiDAR-native decoders alike while keeping the encoder and latent generator fixed. Across KITTI-360, SemanticKITTI, and nuScenes, replacing only the decoder reduces FSVD/FPVD by 50.5% on average across frozen backbones; for generic video VAEs, the reductions reach 71%/74%. On the LiDAR-native LiDM backbone, FRID drops by 71%, with the largest gains at depth discontinuities. In a pretrained LiDM world model, the same zero-shot replacement improves FSVD by 15.5%, narrowing the sim-to-real gap.
Figures & tables
Figure 1: Qualitative comparison across three frozen latent backbones. Blue denotes the ground-truth point cloud, green the inherited VAE decoder, and red our method (CRISP). The three columns show one failure per backbone: SVD, Wan2.1, and LiDM. Across all cases, CRISP suppresses flying pixels and boundary-bridging artifacts that persist in the baseline reconstructions, yielding cleaner geometry and sharper depth contours.
Figure 2: Architecture of CRISP. Left: the frozen encoder Eω maps the input range map y to latent z . Top: the latent adapter Aψ projects and tokenizes z into conditioning tokens τ ; subsequently, the support mask predictor Mϕ takes z and y^d to produce the binary mask m^ . Centre: the cascade DiT decoder Dθ denoises xt on τ at two fusion points (red dashed arrows), yielding the dense depth y^d . Right: m^ and y^d are composed as y^=m^⊙y^d−(1−m^) to produce the final LiDAR range map.
Setting
Dec.
Statistical
Perceptual
JSD ↓
EMD ↓
MMD ↓×10−4
FSVD ↓
FPVD ↓
FRID ↓
Frozen backbones
SVD × 3
base.
0.18961 (00)
0.25550 (00)
2.6946 (44)
156.487 (23)
153.599 (20)
–
ours
0.07145 (02)
0.07653 (01)
2.2505 (20)
40.244 (99)
36.847 (75)
–
Wan × 3
base.
0.16373 (00)
0.11840 (00)
2.2505 (00)
100.917 (00)
113.381 (00)
–
ours
0.07326 (02)
0.08871 (03)
1.7193 (41)
32.534 (40)
31.065 (29)
–
Table 1: Statistical and perceptual reconstruction results. base. uses the original decoder; ours replaces it with CRISP under the same encoder setting. Best values within each backbone pair are bold. Mean over five runs; parenthesized digits give std at mean precision (e.g. 0.123(45)=0.123±0.045 ; (00) denotes std <0.001 at that precision). FRID requires 64-beam range-image features and is reported only for the KITTI-family (LiDM) setting. ↓ / ↑ : lower/higher is better. Same protocol for all tables.
Setting
Dec.
Whole
Edge
Smooth
CD ↓
F ↑
CD ↓
F ↑
CD ↓
F ↑
Frozen backbones
SVD × 3
base.
1.33107 (00)
0.67804 (00)
8.691 (00)
0.0952 (00)
1.18399 (03)
0.73018 (00)
ours
0.63967 (17)
0.80465 (02)
7.364 (00)
0.2383 (00)
0.53709 (09)
0.82158 (02)
Wan × 3
base.
0.77027 (00)
0.73240 (00)
5.861 (00)
0.1038 (00)
0.59749 (00)
0.79303 (00)
ours
0.46453 (09)
0.82551 (01)
5.645 (05)
0.2848 (02)
0.39544 (06)
0.84166 (01)
Table 2: Geometric breakdown into whole-cloud, edge, and smooth sub-clouds. CD ↓ and [email protected]↑ are reported for each region.
Decoder
Statistical
Perceptual
JSD ↓
EMD ↓
MMD ↓×10−4
FSVD ↓
FPVD ↓
FRID ↓
baseline
0.21181 (87)
2.453 (13)
3.980 (44)
37.70 0 (52)
28.85 (34)
137.81 (4.79)
ours
0.21436 (83)
2.420 (13)
3.849 (64)
31.871 (34)
28.11 (18)
132.36 (4.59)
Table 3: Plug-and-play decoding inside a frozen LiDM world model: statistical, perceptual and geometric quality. The latent generator and sampled latents are fixed; only the decoder changes. Values with a decimal point inside std parentheses are written explicitly at the displayed precision.
Figure 3: Qualitative LiDAR reconstruction results (colors as in Fig. 1 ). Columns show, left to right, failure modes CRISP addresses: lost narrow objects, incorrectly filled voids, and flying pixels.
Metric
Inherited
w/o mask
CRISP
Decoder swap
Mask branch
Whole CD ↓
1.331
0.660
0.639
97%
3%
FSVD ↓
156.5
42.6
40.3
98%
2%
JSD ↓
0.190
0.109
0.071
68%
32%
Edge F ↑
0.0952
0.193
0.238
68%
32%
Table 5: Per-component attribution (SVD × 3, nuScenes, converged): share of the Inherited → CRISP gain from the decoder swap alone ( w/o mask ) vs. the mask branch added on top. Inherited from Tables 1 – 2 ; w/o mask and CRISP from Table 14 .
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Removed term
nuScenes val L1
KITTI-family val L1
– (reference)
0.06052
0.0218291
Importance weighting ( w→m in Lvel )
0.07578 ( + 25.2%)
0.0231271 ( + 6.0%)
ℓ1 pixel penalty ( Lx0 )
0.06540 ( + 8.1%)
0.0218987 ( + 0.3%)
Finite differences ( Lder )
0.06291 ( + 3.9%)
0.0235116 ( + 7.7%)
Multi-scale gradient ( Lgrad )
0.06228 ( + 2.9%)
0.0218484 ( + 0.1%)
Appendix
Table 6: Depth-branch leave-one-out: validation L1 ( ↓ ) when each term is removed from the reference objective (for Lvel , its importance weighting is set uniform), on SVD × 3/nuScenes and on the frozen LiDM backbone/KITTI-family; percentages are the change against the matched reference.
Term
λ
Focal BCE ( γ=2 , α=0.4 )
1.0
Soft Dice
1.0
Edge (1st-order)
0.5
Laplacian (2nd-order)
1.5
Hard-pixel mining
0.5
Deep supervision
0.5
Appendix
Table 7: Mask loss sub-term weights (Eq. 10 ).
Removed term
nuScenes Δ IoU
KITTI-family Δ IoU
Laplacian
− 0.00245
− 0.00110
Soft Dice
− 0.00206
− 0.00010
Focal BCE
− 0.00172
− 0.00010
Hard-pixel mining
− 0.00054
≈ 0
Deep supervision
− 0.00034
≈ 0
Edge
− 0.00033
− 0.00020
Appendix
Table 8: Mask-branch leave-one-out: paired change in IoU vs. a matched, same-seed full-objective reference when each term is removed.
Mask objective
IoU
Gain retained
Focal only
0.97418
–
Focal + Dice + Edge
0.98761
91.9%
Focal + Dice + Laplacian
0.98831
96.7%
Full six-term
0.98879
100%
Appendix
Table 9: Mask-branch minimal subset (LiDM backbone, KITTI-family, 10 epochs): converged IoU when zeroing all but the listed terms, and the resulting share of the full objective’s IoU gain over focal-only supervision retained.
Term
Failure mode
Provenance
Weight
Reduced
Depth branch (Eq. 7 )
Lvel
region: boundary-aware velocity regression on valid pixels
Table 10: Loss-term justification: failure mode, operator provenance, weight, and reduced-objective status for every term in Ldepth (Eq. 7 ) and Lmask (Eq. 10 ).
Objective
Branches
Terms
ms/step ↓
Peak GiB ↓
Depth only ( Lvel+Lvel,raw )
Depth
2
165.11
19.878
+Lfocal
Depth + Mask
3
263.23
34.397
Reduced set
Depth + Mask
7
265.68
34.393
Full objective
Depth + Mask
11
271.04
34.394
Appendix
Table 11: Training step time and peak memory as loss terms are added, network/data/checkpoints held fixed (one H200, BF16, SVD × 3/nuScenes, median of 100 steps).
Table 12: Per-frame CD ( ↓ ) and [email protected] ( ↑ ) for Wan × 3 (left) and Wan × 1 (right) under joint temporal encoding. Frame t=0 is systematically the best; the degradation motivates per-frame inference. Bold = best per column.
Loss / setting
Frozen encoder ( × 3)
Encoder-adapted ( × 1)
F1
F2
A1
A2
A3
Encoder trainable
no
no
yes
yes
no
Mask branch
–
yes
–
yes
yes
Velocity loss
L 2
L 1
L 2
L 2
L 1
λx0
0.0
0.5
0.0
0.0
0.5
λgrad
0.0
0.5
0.0
0.0
0.5
Appendix
Table 13: Training stages and loss weights. Frozen-encoder variants ( × 3) use Stages F1-F2; encoder-adapted variants ( × 1) use Stages A1-A3. “–” denotes an inactive branch or loss.
Figure 4: Step ablation on nuScenes with the SVD × 3 backbone ( n=33,021 LiDAR frames). Each panel plots a different metric group as a function of Euler sampling steps; the dashed vertical line marks the adopted 5-step setting. CD and perceptual metrics exhibit opposing trends, motivating a deliberate operating-point choice rather than optimising either metric alone.
Variant
Statistical
Perceptual
JSD ↓
EMD ↓
MMD ↓×10−4
FSVD ↓
FPVD ↓
FRID ↓
GT mask, 35 epochs
w/o early fusion
0.120
0.241
3.101
119.557
100.839
-
w/ early fusion
0.104
0.196
2.739
98.649
81.651
-
Full system, converged
Loss everywhere (no mask branch)
0.109
0.110
2.371
42.556
35.993
-
Appendix
Table 14: DiT decoder ablation (SVD × 3, nuScenes). Upper : GT support mask applied, 35 epochs (3-run mean); early-fusion comparison. Lower : converged end-to-end; mask-formulation comparison. Best within each group is bold.
Config
IoU ↑
Prec. ↑
Rec. ↑
F1 ↑
Loss composition (depth + latent, 10 ep.)
Focal only
0.758
0.860
0.865
0.862
+ Dice
0.827
0.864
0.951
0.906
+ Edge loss
0.838
0.871
0.957
0.912
+ Laplacian
0.841
0.869
0.963
0.914
+ Pixel mining
0.845
0.874
0.962
0.916
Appendix
Table 15: Mask predictor ablation (SVD × 3, nuScenes). Upper : cumulative loss composition, depth + latent input, 10 ep. each. Lower : input conditioning, full loss set, converged. † Stage 6 is the deployed configuration; the marginal gap vs. Stage 5 is within single-run noise but deep supervision is retained for robustness on thin-structure regions.
Design axis
Variant
Val L1 ↓
Latent conditioning
Cross-attention
0.1872
Register tokens
0.1156
Early+mid concat (ours)
0.0446
Euler steps
1
0.0643
3
0.0527
5 (ours)
0.0446
Appendix
Table 16: Design-choice replication on the frozen LiDAR-native LiDM backbone (KITTI-family). Validation L1 ( ↓ ); cf. the SVD × 3/nuScenes ablations in Table 14 and Fig. 4 .
Backbone
Encoder (ms/f)
Decoder (ms/f)
End-to-end
Base.
CRISP
Base. fps
CRISP fps
LiDM
1.31
1.39
23.06
369
41
SVD × 3
0.74
1.49
0 3.92
448
215
Wan × 3
7.78
9.71
0 3.84 †
0 57
86
Appendix
Table 17: End-to-end inference throughput (ms/frame and fps). Encoder cost is shared between baseline and CRISP. Measured on 8 × H200, torch.compile , batch 25, 5 Euler steps. † CRISP decoder is faster than the inherited Wan2.1 VAE decoder.
Asset
Reference
License
KITTI-360
[ 28 ]
CC BY-NC-SA 3.0
SemanticKITTI
[ 2 ]
CC BY-NC-SA 4.0
nuScenes
[ 5 ]
CC BY-NC-SA 4.0
LiDM
[ 42 ]
MIT
SVD VAE
[ 3 ]
Stability AI Community License
Wan2.1 VAE
[ 53 ]
Apache 2.0
Appendix
Table 18: Licenses for datasets, pretrained models, and third-party code used in this work.
Figure 5: Extended qualitative comparison, SVD × 3 (frozen encoder, nuScenes). The baseline decoder introduces heavy scan-line scatter at depth discontinuities: beam-level structure of a vehicle and pole, visible in ground truth, dissolves into noise clouds in the baseline reconstruction. CRISP restores continuous scan-line geometry and suppresses the boundary-bridging artifacts highlighted by the inset box.
Figure 6: Extended qualitative comparison, SVD × 1 (encoder-adapted, nuScenes). The baseline decoder misplaces a mid-range object cluster (circled tree trunks): points smear outward rather than forming a compact boundary. CRISP localises the cluster accurately and sharpens the depth contour at the foreground-background transition indicated by the inset box.
Figure 7: Extended qualitative comparison, Wan × 3 (frozen encoder, nuScenes). The highlighted region contains a foreground-background scene depth gap that the baseline decoder bridges: points fill in between the two surfaces, erasing the geometric discontinuity. CRISP retains the gap as a sharp void, consistent with the ground-truth scan.
Figure 8: Extended qualitative comparison, Wan × 1 (encoder-adapted, nuScenes). The baseline decoder incorrectly fills the open void region visible in ground truth (inset box): spurious returns appear where no surface exists. CRISP preserves the void, producing a sparse valid-return map that matches the support structure of the real scan.
Figure 9: Extended qualitative comparison, LiDM (frozen encoder, KITTI family). The 64-beam KITTI geometry makes foreground-background depth gaps particularly sharp. The baseline decoder bridges the gap in the highlighted region, collapsing two geometrically distinct surfaces. CRISP reproduces the correct depth discontinuity, matching the gap geometry visible in the ground-truth scan.
Figure 10: Extended qualitative comparison, LiDM world model (plug-and-play decoding, KITTI family). No ground truth is shown because these are generative samples from the frozen LiDM latent generator; both rows decode the same sampled latent. The baseline decoder introduces grid-aligned blocky artifacts and smears foreground boundaries. CRISP, swapped in without retraining the generator, produces cleaner scan geometry and sharper depth contours from the same latent input.