cs.CVDec 19, 2025

Event-based Scene Synthesis via Inter-Frame Residual Alignment

Authors: Jiyun Kong, Jun-Hyuk Kim, Jong-Seok Lee

Organizations: Yonsei University, Korea · Chung-Ang University, Korea

Abstract

Event-based scene synthesis reconstructs target RGB frames from sparse image observations and asynchronous event streams, encompassing both video frame prediction and interpolation. Existing event-based synthesis methods commonly estimate optical flow to warp the observed frames toward the target time, but are vulnerable to inaccurate flow under large motion and occlusion and often rely on flow supervision or pretrained estimators. In this work, we propose EvFRA, an Event-based scene synthesis framework based on inter-Frame Residual Alignment. We identify a structural correspondence between event measurements and frame-to-frame scene changes, and exploit this correspondence for target frame synthesis. Our training pipeline consists of two stages: 1) an Event-to-Residual Alignment Variational Autoencoder (ER-VAE) aligns the event frame captured between the anchor and target frames with the corresponding inter-frame residual, and 2) a ControlNet-conditioned diffusion model is fine-tuned to denoise the residual latent using event data. Our method outperforms state-of-the-art methods by up to 2.61 dB and 1.85 dB in PSNR for frame prediction and interpolation, respectively, with consistent SSIM improvements. Code is available at https://github.com/jiyun-kong/EvFRA.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. UniE2F: A Unified Diffusion Framework for Event-to-Frame Reconstruction with Video Foundation Models

    Feb 22, 2026Gang Xu, Zhiyu Zhu, Junhui HouVideo Diffusion ModelsLarge Inter-Frame Displacements

  2. LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models

    Jul 9, 2026Cheng-De Fan, Chun-Wei Tuan Mu, Chen-Wei Chang +4Video Diffusion ModelsFuture Video Prediction

  3. Bridging Event Streams and DiT: Event-Guided Video Frame Interpolation

    Aug 11, 2026Guixu Lin, Yuyang Yu, Xiang Ji +6Large Inter-Frame DisplacementsDiffusion Transformers