CRISP: Fixing Flying Pixels in Latent LiDAR Generation via Diffusion Decoding
Organizations: TU München · BMW AG
Abstract
Latent LiDAR pipelines suffer from flying pixels: convolutional VAEs blur sharp radial depth discontinuities, yielding edge depths that back-project to points floating between surfaces. We identify this as a major, directly correctable decoder bottleneck and introduce CRISP: a pixel-space diffusion decoder with a backbone-agnostic latent adapter, DiT-based denoiser, and support mask predictor. CRISP replaces video-VAE and LiDAR-native decoders alike while keeping the encoder and latent generator fixed. Across KITTI-360, SemanticKITTI, and nuScenes, replacing only the decoder reduces FSVD/FPVD by 50.5% on average across frozen backbones; for generic video VAEs, the reductions reach 71%/74%. On the LiDAR-native LiDM backbone, FRID drops by 71%, with the largest gains at depth discontinuities. In a pretrained LiDM world model, the same zero-shot replacement improves FSVD by 15.5%, narrowing the sim-to-real gap.
Figures & tables
| Setting | Dec. | Statistical | Perceptual | ||||
| JSD | EMD | MMD | FSVD | FPVD | FRID | ||
| Frozen backbones | |||||||
| SVD 3 | base. | 0.18961 (00) | 0.25550 (00) | 2.6946 (44) | 156.487 (23) | 153.599 (20) | – |
| ours | 0.07145 (02) | 0.07653 (01) | 2.2505 (20) | 40.244 (99) | 36.847 (75) | – | |
| Wan 3 | base. | 0.16373 (00) | 0.11840 (00) | 2.2505 (00) | 100.917 (00) | 113.381 (00) | – |
| ours | 0.07326 (02) | 0.08871 (03) | 1.7193 (41) | 32.534 (40) | 31.065 (29) | – | |
| Setting | Dec. | Whole | Edge | Smooth | |||
| CD | F | CD | F | CD | F | ||
| Frozen backbones | |||||||
| SVD 3 | base. | 1.33107 (00) | 0.67804 (00) | 8.691 (00) | 0.0952 (00) | 1.18399 (03) | 0.73018 (00) |
| ours | 0.63967 (17) | 0.80465 (02) | 7.364 (00) | 0.2383 (00) | 0.53709 (09) | 0.82158 (02) | |
| Wan 3 | base. | 0.77027 (00) | 0.73240 (00) | 5.861 (00) | 0.1038 (00) | 0.59749 (00) | 0.79303 (00) |
| ours | 0.46453 (09) | 0.82551 (01) | 5.645 (05) | 0.2848 (02) | 0.39544 (06) | 0.84166 (01) | |
| Decoder | Statistical | Perceptual | ||||
| JSD | EMD | MMD | FSVD | FPVD | FRID | |
| baseline | 0.21181 (87) | 2.453 (13) | 3.980 (44) | 37.70 0 (52) | 28.85 (34) | 137.81 (4.79) |
| ours | 0.21436 (83) | 2.420 (13) | 3.849 (64) | 31.871 (34) | 28.11 (18) | 132.36 (4.59) |
| Backbone | Decoder params | GFLOPs | Decoder ms/frame | ||
| Inh. CRISP | Ratio | Inh. CRISP | Ratio | (Inh. CRISP, ratio) | |
| SVD 3 | 63.6M 516.0M | 8.1 | 376.6 1290.7 | 3.4 | 1.49 3.92 (2.63 ) |
| Wan 3 | 73.3M 516.1M | 7.0 | 352.6 1291.4 | 3.7 | 9.71 3.84 (0.40 ) |
| LiDM | 8.6M 516.0M | 60.0 | 119.0 2717.2 | 22.8 | 1.39 23.06 (16.6 ) |
| Metric | Inherited | w/o mask | CRISP | Decoder swap | Mask branch |
| Whole CD | 1.331 | 0.660 | 0.639 | 97% | 3% |
| FSVD | 156.5 | 42.6 | 40.3 | 98% | 2% |
| JSD | 0.190 | 0.109 | 0.071 | 68% | 32% |
| Edge F | 0.0952 | 0.193 | 0.238 | 68% | 32% |
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| Removed term | nuScenes val L1 | KITTI-family val L1 |
| – (reference) | 0.06052 | 0.0218291 |
| Importance weighting ( in ) | 0.07578 ( 25.2%) | 0.0231271 ( 6.0%) |
| pixel penalty ( ) | 0.06540 ( 8.1%) | 0.0218987 ( 0.3%) |
| Finite differences ( ) | 0.06291 ( 3.9%) | 0.0235116 ( 7.7%) |
| Multi-scale gradient ( ) | 0.06228 ( 2.9%) | 0.0218484 ( 0.1%) |
| Term | |
| Focal BCE ( , ) | 1.0 |
| Soft Dice | 1.0 |
| Edge (1st-order) | 0.5 |
| Laplacian (2nd-order) | 1.5 |
| Hard-pixel mining | 0.5 |
| Deep supervision | 0.5 |
| Removed term | nuScenes IoU | KITTI-family IoU |
| Laplacian | 0.00245 | 0.00110 |
| Soft Dice | 0.00206 | 0.00010 |
| Focal BCE | 0.00172 | 0.00010 |
| Hard-pixel mining | 0.00054 | 0 |
| Deep supervision | 0.00034 | 0 |
| Edge | 0.00033 | 0.00020 |
| Mask objective | IoU | Gain retained |
| Focal only | 0.97418 | – |
| Focal + Dice + Edge | 0.98761 | 91.9% |
| Focal + Dice + Laplacian | 0.98831 | 96.7% |
| Full six-term | 0.98879 | 100% |
| Term | Failure mode | Provenance | Weight | Reduced |
| Depth branch (Eq. 7 ) | ||||
| region: boundary-aware velocity regression on valid pixels | – | 1.0 | yes | |
| stable unweighted baseline; normalises ’s per-edge gradient share | structural, no operator citation | 1.0 | yes | |
| anisotropic per-ring radial bias invisible per-pixel | per-axis derivative matching (cf. LiDM gradient loss [ 42 ] ) | Table 13 | yes | |
| low-noise accuracy of (fine-tuning signal) | timestep-weighted diffusion loss [ 45 , 48 ] | 0.5 | yes | |
| low-frequency range-gradient residual | LiDM multi-scale gradient loss [ 42 ] | 0.5 | no | |
| Objective | Branches | Terms | ms/step | Peak GiB |
| Depth only ( ) | Depth | 2 | 165.11 | 19.878 |
| Depth + Mask | 3 | 263.23 | 34.397 | |
| Reduced set | Depth + Mask | 7 | 265.68 | 34.393 |
| Full objective | Depth + Mask | 11 | 271.04 | 34.394 |
| Wan 3 | Wan 1 | |||
| CD | [email protected] | CD | [email protected] | |
| 0 | 0.770 | 0.732 | 0.309 | 0.867 |
| 1 | 0.792 | 0.718 | 0.336 | 0.862 |
| 2 | 0.805 | 0.712 | 0.350 | 0.859 |
| 3 | 0.794 | 0.721 | 0.333 | 0.862 |
| 4 | 0.804 | 0.716 | 0.351 | 0.858 |
| Loss / setting | Frozen encoder ( 3) | Encoder-adapted ( 1) | |||
| F1 | F2 | A1 | A2 | A3 | |
| Encoder trainable | no | no | yes | yes | no |
| Mask branch | – | yes | – | yes | yes |
| Velocity loss | L 2 | L 1 | L 2 | L 2 | L 1 |
| 0.0 | 0.5 | 0.0 | 0.0 | 0.5 | |
| 0.0 | 0.5 | 0.0 | 0.0 | 0.5 | |
| Variant | Statistical | Perceptual | ||||
| JSD | EMD | MMD | FSVD | FPVD | FRID | |
| GT mask, 35 epochs | ||||||
| w/o early fusion | 0.120 | 0.241 | 3.101 | 119.557 | 100.839 | - |
| w/ early fusion | 0.104 | 0.196 | 2.739 | 98.649 | 81.651 | - |
| Full system, converged | ||||||
| Loss everywhere (no mask branch) | 0.109 | 0.110 | 2.371 | 42.556 | 35.993 | - |
| Config | IoU | Prec. | Rec. | F1 |
| Loss composition (depth + latent, 10 ep.) | ||||
| Focal only | 0.758 | 0.860 | 0.865 | 0.862 |
| Dice | 0.827 | 0.864 | 0.951 | 0.906 |
| Edge loss | 0.838 | 0.871 | 0.957 | 0.912 |
| Laplacian | 0.841 | 0.869 | 0.963 | 0.914 |
| Pixel mining | 0.845 | 0.874 | 0.962 | 0.916 |
| Design axis | Variant | Val L1 |
| Latent conditioning | Cross-attention | 0.1872 |
| Register tokens | 0.1156 | |
| Early+mid concat (ours) | 0.0446 | |
| Euler steps | 1 | 0.0643 |
| 3 | 0.0527 | |
| 5 (ours) | 0.0446 |
| Backbone | Encoder (ms/f) | Decoder (ms/f) | End-to-end | ||
| Base. | CRISP | Base. fps | CRISP fps | ||
| LiDM | 1.31 | 1.39 | 23.06 | 369 | 41 |
| SVD 3 | 0.74 | 1.49 | 0 3.92 | 448 | 215 |
| Wan 3 | 7.78 | 9.71 | 0 3.84 † | 0 57 | 86 |
| Asset | Reference | License |
| KITTI-360 | [ 28 ] | CC BY-NC-SA 3.0 |
| SemanticKITTI | [ 2 ] | CC BY-NC-SA 4.0 |
| nuScenes | [ 5 ] | CC BY-NC-SA 4.0 |
| LiDM | [ 42 ] | MIT |
| SVD VAE | [ 3 ] | Stability AI Community License |
| Wan2.1 VAE | [ 53 ] | Apache 2.0 |