Event-based Scene Synthesis via Inter-Frame Residual Alignment
Organizations: Yonsei University, Korea · Chung-Ang University, Korea
Abstract
Event-based scene synthesis reconstructs target RGB frames from sparse image observations and asynchronous event streams, encompassing both video frame prediction and interpolation. Existing event-based synthesis methods commonly estimate optical flow to warp the observed frames toward the target time, but are vulnerable to inaccurate flow under large motion and occlusion and often rely on flow supervision or pretrained estimators. In this work, we propose EvFRA, an Event-based scene synthesis framework based on inter-Frame Residual Alignment. We identify a structural correspondence between event measurements and frame-to-frame scene changes, and exploit this correspondence for target frame synthesis. Our training pipeline consists of two stages: 1) an Event-to-Residual Alignment Variational Autoencoder (ER-VAE) aligns the event frame captured between the anchor and target frames with the corresponding inter-frame residual, and 2) a ControlNet-conditioned diffusion model is fine-tuned to denoise the residual latent using event data. Our method outperforms state-of-the-art methods by up to 2.61 dB and 1.85 dB in PSNR for frame prediction and interpolation, respectively, with consistent SSIM improvements. Code is available at https://github.com/jiyun-kong/EvFRA.
Figures & tables
| Task | Method | BS-ERGB [ 47 ] | HS-ERGB [ 48 ] | |||||||
| 1 frame | 3 frames | 7 frames | ||||||||
| PSNR | SSIM | LPIPS | PSNR | SSIM | LPIPS | PSNR | SSIM | LPIPS | ||
| Frame Prediction | VFPSIE [ 59 ] (AAAI’24) | 24.41 | 0.805 | 0.091 | 23.20 | 0.782 | 0.114 | 28.81 | 0.845 | 0.112 |
| EvFRA(VFP) | 25.27 | 0.895 | 0.100 | 24.20 | 0.804 | 0.119 | 29.74 | 0.858 | 0.101 | |
| Frame Interpolation | TimeLens [ 48 ] (CVPR’21) | 23.08 | 0.694 | 0.121 | 21.93 | 0.658 | 0.172 | 25.98 | 0.793 | 0.167 |
| CBMNet-Large [ 25 ] (CVPR’23) | 24.96 | 0.827 | 0.120 | 23.83 | 0.800 | 0.123 | 28.14 | 0.844 | 0.149 | |
| Task | Method | GoPro [ 31 ] | |||||
| 7 frames | 15 frames | ||||||
| PSNR | SSIM | LPIPS | PSNR | SSIM | LPIPS | ||
| Frame Prediction | VFPSIE [ 59 ] (AAAI’24) | 18.26 | 0.557 | 0.386 | 17.53 | 0.523 | 0.438 |
| EvFRA(VFP) | 20.87 | 0.703 | 0.231 | 19.75 | 0.672 | 0.308 | |
| Frame Interpolation | TimeLens [ 48 ] (CVPR’21) | 17.29 | 0.571 | 0.300 | 15.02 | 0.531 | 0.384 |
| CBMNet-Large [ 25 ] (CVPR’23) | 17.98 | 0.592 | 0.327 | 17.44 | 0.532 | 0.372 | |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Metrics | Seed #123 | Seed #0 | Seed #42 | Mean Std |
| Video Frame Prediction | PSNR | 24.20 | 23.92 | 24.09 | 24.07 0.14 |
| SSIM | 0.804 | 0.800 | 0.802 | 0.802 0.002 | |
| LPIPS | 0.119 | 0.119 | 0.120 | 0.119 0.001 | |
| Video Frame Interpolation | PSNR | 24.28 | 24.23 | 24.28 | 24.26 0.03 |
| SSIM | 0.831 | 0.827 | 0.827 | 0.828 0.002 | |
| LPIPS | 0.116 | 0.115 | 0.115 | 0.115 0.001 |
| ID | Target | Initialization | Event Cond. | Generator | PSNR | SSIM | LPIPS |
| A | ✓ | Diffusion | 16.72 | 0.571 | 0.340 | ||
| B | ✓ | Diffusion | 19.51 | 0.703 | 0.315 | ||
| C | ✓ | Diffusion | 19.96 | 0.711 | 0.313 | ||
| D | ✓ | Diffusion | 23.67 | 0.762 | 0.161 | ||
| E | ✓ | Diffusion | 21.43 | 0.734 | 0.204 | ||
| F (EvFRA) | ✓ | Diffusion | 24.20 | 0.804 | 0.119 |
| Steps | 2 | 5 (default) | 10 | 15 | 25 | 50 |
| PSNR | 24.15 | 24.20 | 24.20 | 24.21 | 24.21 | 24.21 |
| SSIM | 0.800 | 0.804 | 0.802 | 0.801 | 0.801 | 0.801 |
| LPIPS | 0.119 | 0.119 | 0.120 | 0.120 | 0.120 | 0.120 |
| Time (s) | 7.1 | 10.3 | 17.7 | 22.9 | 35.4 | 67.5 |
| Order | PSNR | SSIM | LPIPS |
| 25.27 / 24.20 | 0.895 / 0.804 | 0.100 / 0.119 | |
| 22.58 / 22.07 | 0.775 / 0.735 | 0.150 / 0.181 | |
| 21.39 / 20.79 | 0.755 / 0.701 | 0.159 / 0.179 |
| Setting | 1 frame | 3 frames | ||||
| PSNR | SSIM | LPIPS | PSNR | SSIM | LPIPS | |
| VAE Scaling | ||||||
| w/o | 25.20 | 0.882 | 0.100 | 24.13 | 0.801 | 0.119 |
| with | 25.27 | 0.895 | 0.100 | 24.20 | 0.804 | 0.119 |
| KL Divergence | ||||||
| w/o | 25.27 | 0.895 | 0.100 | 24.20 | 0.804 | 0.119 |
| Metric | |||||
| PSNR | 22.51 | 23.01 | 23.02 | 22.47 | 22.42 |
| SSIM | 0.725 | 0.746 | 0.738 | 0.735 | 0.698 |
| LPIPS | 0.182 | 0.160 | 0.169 | 0.168 | 0.178 |
| Modality | Task | Method | BS-ERGB [ 47 ] | HS-ERGB [ 48 ] | |||||||
| 1 frame | 3 frames | 7 frames | |||||||||
| PSNR | SSIM | LPIPS | PSNR | SSIM | LPIPS | PSNR | SSIM | LPIPS | |||
| 2 Images | Frame Prediction | RaMViD [ 17 ] (TMLR’22) | 14.40 | 0.451 | 0.400 | 14.32 | 0.450 | 0.392 | 19.28 | 0.691 | 0.293 |
| DMVFN [ 18 ] (CVPR’23) | 23.67 | 0.675 | 0.093 | 22.08 | 0.652 | 0.125 | 28.14 | 0.788 | 0.122 | ||
| Event | Reconstruction | E2HQV [ 34 ] (AAAI’24) | 14.19 | 0.358 | 0.247 | 13.88 | 0.358 | 0.260 | 17.91 | 0.487 | 0.329 |
| HyperE2VID [ 9 ] (TIP’24) | 18.79 | 0.487 | 0.273 | 18.78 | 0.410 | 0.278 | 18.20 | 0.529 | 0.331 | ||
| Modality | Task | Method | GoPro [ 31 ] | |||||
| 7 frames | 15 frames | |||||||
| PSNR | SSIM | LPIPS | PSNR | SSIM | LPIPS | |||
| 2 Images | Frame Prediction | RaMViD [ 17 ] (TMLR’22) | 13.50 | 0.487 | 0.457 | 13.32 | 0.484 | 0.560 |
| DMVFN [ 18 ] (CVPR’23) | 14.65 | 0.442 | 0.451 | 13.12 | 0.413 | 0.559 | ||
| Event | Reconstruction | E2HQV [ 34 ] (AAAI’24) | 10.81 | 0.468 | 0.423 | 10.86 | 0.471 | 0.520 |
| HyperE2VID [ 9 ] (TIP’24) | 9.72 | 0.430 | 0.468 | 9.62 | 0.430 | 0.571 | ||
| Stage | Small | Medium | Large | Occlusion |
| VFPSIE Pred. Flow Warp | 30.24 | 27.54 | 20.95 | 23.63 |
| RAFT Flow Warp | 31.26 | 29.98 | 25.03 | 25.01 |
| VFPSIE Final | 30.55 | 28.47 | 22.68 | 25.34 |
| EvFRA | 33.87 | 31.85 | 27.63 | 29.17 |
| Task | Method | GFLOPs | # Params (M) | Run Time (s) | Memory Usage (GB) | Notes |
| Frame Prediction | VFPSIE [ 59 ] | 55.3 | 2.15 | 0.0147 | 0.457 | Optical flow-based |
| EvFRA(VFP) | 10275.0 | 2029.76 | 10.3 | 4.310 | Diffusion Model | |
| Frame Interpolation | TimeLens [ 48 ] | 2456.4 | 79.21 | 0.133 | 1.73 | Optical flow-based |
| CBMNet-Large [ 25 ] | 3201.1 | 22.23 | 1.163 | 10.161 | Optical flow-based | |
| RE-VDM [ 6 ] | 23921.9 | 2936.70 | 82.296 | 9.388 | Video Diffusion Model | |
| EvFRA(VFI) | 20953.1 | 2029.76 | 21.46 | 4.310 | Diffusion Model |
| Metric | Win Count | Mean Improvement | Wilcoxon | -test |
| PSNR | 15 / 15 | 0.930 | 0.00051 | 0.00238 |
| SSIM | 13 / 15 | 0.014 | 0.00098 | 0.00047 |
| LPIPS | 13 / 15 | 0.011 | 0.00092 | 0.00044 |
| Example | Ground Truth | Prediction |
| Basketball | 0.968 | 0.994 |
| Football | 0.965 | 0.996 |