cs.CVOct 8, 2026

CRISP: Fixing Flying Pixels in Latent LiDAR Generation via Diffusion Decoding

Authors: Andrea Ceron, Michael Schmidt, Alvaro Marcos-Ramiro, Sebastian Schmidt, Benjamin Busam

Organizations: TU München · BMW AG

Abstract

Latent LiDAR pipelines suffer from flying pixels: convolutional VAEs blur sharp radial depth discontinuities, yielding edge depths that back-project to points floating between surfaces. We identify this as a major, directly correctable decoder bottleneck and introduce CRISP: a pixel-space diffusion decoder with a backbone-agnostic latent adapter, DiT-based denoiser, and support mask predictor. CRISP replaces video-VAE and LiDAR-native decoders alike while keeping the encoder and latent generator fixed. Across KITTI-360, SemanticKITTI, and nuScenes, replacing only the decoder reduces FSVD/FPVD by 50.5% on average across frozen backbones; for generic video VAEs, the reductions reach 71%/74%. On the LiDAR-native LiDM backbone, FRID drops by 71%, with the largest gains at depth discontinuities. In a pretrained LiDM world model, the same zero-shot replacement improves FSVD by 15.5%, narrowing the sim-to-real gap.

Figures & tables

Appendix figures & tables20 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. PointDiffusion: Diffusion-Based Scene Completion in the Point Cloud Domain

    Jun 14, 2026Chidera Agbasiere, Mikhail Sannikov, Faith Ogunwoye +5Point Cloud Completion3D Reconstruction

  2. LiDAR Resolution Recovery via Foundation-Model-Guided Diffusion

    Oct 6, 2026Samed Doğan, Nico Leuze, Alfred SchöttlDiffusion Model GuidanceDepth Completion