PickMoment: Continuous-Time Single-Image-to-Video via Learning Deblurring and Blur-to-Video
Organizations: Department of Artificial Intelligence, Hanyang University · Department of Intelligence Convergence, Hanyang University · Department of Computer Science, Hanyang University
Abstract
Motion blur arises from the temporal integration of a continuous sharp signal over a finite exposure window, yet existing learning-based methods sidestep this physical model and predict only the sharp signal itself: most single-image deblurring methods recover a single frame at the exposure center, while blur-to-video methods predict a fixed set of frames. We introduce PickMoment, a continuous-time reformulation that directly learns the interval-mean blur over arbitrary sub-intervals of the exposure with a single deterministic model. Drawing an analogy to MeanFlow's average-velocity formulation, we train the model with three supervisions derived from the blur integral: an empirical reconstruction loss from available subframes, an additivity loss that enforces self-consistency across overlapping sub-intervals, and a sharp-frame loss anchored at the zero-interval limit. A single trained model unifies single-image deblurring, blur-to-video generation, and continuous-time pick-a-moment recovery as different queries to the same network, with no separate training for each task. Our PickMoment achieves state-of-the-art performance among generative-based deblurring methods on GoPro and HIDE while competitive against restoration-based methods on RealBlur, and the highest per-frame fidelity on GoPro-7 blur-to-video, all in a single forward pass without iterative sampling.
Figures & tables
| Restoration-based | Generative-based | |||||||||
| Dataset | Metric | Restormer [ 45 ] | HI-Diff [ 5 ] | UFPNet [ 8 ] | AdaRevD [ 27 ] | DiffBIR [ 20 ] | OSEDiff [ 43 ] | Diff-Plugin [ 24 ] | FideDiff [ 22 ] | Ours |
| GoPro | PSNR | 32.92 | 33.33 | 34.09 | 34.60 | 26.15 | 24.34 | 22.88 | 28.79 | 29.43 |
| SSIM | 0.9611 | 0.9642 | 0.9686 | 0.9716 | 0.8377 | 0.8228 | 0.7798 | 0.9148 | 0.9253 | |
| LPIPS | 0.0841 | 0.0799 | 0.0764 | 0.0712 | 0.2366 | 0.1738 | 0.2332 | 0.0831 | 0.0802 | |
| DISTS | 0.0724 | 0.0710 | 0.0666 | 0.0672 | 0.1460 | 0.0834 | 0.1166 | 0.0525 | 0.0429 | |
| HIDE | PSNR | 31.22 | 31.46 | 31.74 | 32.35 | 25.94 | 23.20 | 21.94 | 27.28 | 27.52 |
| Deblurring | Blur-to-video | ||||||||
| Method | PSNR | SSIM | LPIPS | DISTS | PSNR p | SSIM p | LPIPS p | FVD | EPE |
| 27.72 | 0.8944 | 0.1595 | 0.1149 | 29.16 | 0.8894 | 0.0292 | 59.51 | 1.15 | |
| 27.76 | 0.8945 | 0.1604 | 0.1166 | 29.26 | 0.8933 | 0.0278 | 58.76 | 1.13 | |
| (ours) | 28.45 | 0.9080 | 0.1455 | 0.1043 | 29.78 | 0.9034 | 0.0271 | 59.29 | 1.02 |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Backbone | FLUX.2-klein-4B (transformer only fine-tuned, VAE frozen) |
| Trainable params | 4B (full transformer) |
| Optimizer | Adam |
| Total iterations | 100K |
| LR schedule | Linear warmup to over 2K, then cosine-annealing to |
| Mixed precision | bf16 |
| Gradient accumulation | 2 steps |
| PSNR (dB) | Description | Mean | Median | Std. | 10th pct. | 90th pct. |
| vs. | (VAE round-trip only) | 42.70 | 43.12 | 2.32 | 39.27 | 45.24 |
| vs. | (pure additivity) | 39.82 | 39.38 | 4.49 | 34.36 | 45.94 |
| vs. | (combined) | 37.40 | 37.26 | 2.80 | 33.83 | 41.03 |
| Method | PSNR | SSIM | LPIPS | DISTS |
| Jin et al. [ 15 ] | 29.44 | 0.9237 | 0.1345 | 0.0982 |
| MotionETR [ 48 ] | 28.60 | 0.9415 | 0.0509 | 0.0495 |
| Blur2Vid [ 37 ] | 30.66 | 0.9528 | 0.0363 | 0.0380 |
| Ours | 33.22 | 0.9664 | 0.0291 | 0.0298 |
| Dataset | Metric | Full | |||
| GoPro | PSNR | 28.45 | 28.56 | 29.24 | 29.43 |
| SSIM | 0.9080 | 0.9096 | 0.9236 | 0.9253 | |
| LPIPS | 0.1455 | 0.1422 | 0.0787 | 0.0802 | |
| DISTS | 0.1043 | 0.1020 | 0.0420 | 0.0429 | |
| HIDE | PSNR | 26.73 | 26.79 | 27.49 | 27.52 |
| SSIM | 0.8668 | 0.8687 | 0.8894 | 0.8891 |
| Asset | Type | Citation | License |
| GoPro_Large_All | Dataset | [ 28 ] | Research use (per authors) |
| HIDE | Dataset | [ 35 ] | Research use (per authors) |
| RealBlur-J / RealBlur-R | Dataset | [ 32 ] | CC BY-NC-SA 4.0 |
| FLUX.2-klein | Pretrained model | [ 3 ] | FLUX.2 Community License (non-commercial) |
| Stable Diffusion (referenced) | Pretrained model | [ 33 ] | CreativeML Open RAIL-M |
| RAFT | Optical flow | [ 38 ] | BSD 3-Clause |