PickMoment: Continuous-Time Single-Image-to-Video via Learning Deblurring and Blur-to-Video
Authors: Junseong Shin, Hyeonsu Jo, Daehyun Kim, Tae Hyun Kim
Organizations: Department of Artificial Intelligence, Hanyang University · Department of Intelligence Convergence, Hanyang University · Department of Computer Science, Hanyang University
Motion blur arises from the temporal integration of a continuous sharp signal over a finite exposure window, yet existing learning-based methods sidestep this physical model and predict only the sharp signal itself: most single-image deblurring methods recover a single frame at the exposure center, while blur-to-video methods predict a fixed set of frames. We introduce PickMoment, a continuous-time reformulation that directly learns the interval-mean blur over arbitrary sub-intervals of the exposure with a single deterministic model. Drawing an analogy to MeanFlow's average-velocity formulation, we train the model with three supervisions derived from the blur integral: an empirical reconstruction loss from available subframes, an additivity loss that enforces self-consistency across overlapping sub-intervals, and a sharp-frame loss anchored at the zero-interval limit. A single trained model unifies single-image deblurring, blur-to-video generation, and continuous-time pick-a-moment recovery as different queries to the same network, with no separate training for each task. Our PickMoment achieves state-of-the-art performance among generative-based deblurring methods on GoPro and HIDE while competitive against restoration-based methods on RealBlur, and the highest per-frame fidelity on GoPro-7 blur-to-video, all in a single forward pass without iterative sampling.
Figures & tables
Figure 1 : Overview of the proposed PickMoment framework. Given a single motion-blurred image, our method unifies three tasks: (a) Image deblurring, predicting the center sharp frame of the blur exposure. (b) Blur-to-video generation predicts an n -frame sharp video by uniformly sampling the normalized exposure trajectory at M time points with interval Δt=1/(M−1) ; the figure illustrates an 11-frame example. (c) Arbitrary sharp-frame prediction enables the recovery of sharp frames at arbitrary time points within the blur exposure, thereby enabling continuous-time video generation.
Figure 2 : Conceptual illustration of the training strategy of PickMoment.
Restoration-based
Generative-based
Dataset
Metric
Restormer [ 45 ]
HI-Diff [ 5 ]
UFPNet [ 8 ]
AdaRevD [ 27 ]
DiffBIR [ 20 ]
OSEDiff [ 43 ]
Diff-Plugin [ 24 ]
FideDiff [ 22 ]
Ours
GoPro
PSNR ↑
32.92
33.33
34.09
34.60
26.15
24.34
22.88
28.79
29.43
SSIM ↑
0.9611
0.9642
0.9686
0.9716
0.8377
0.8228
0.7798
0.9148
0.9253
LPIPS ↓
0.0841
0.0799
0.0764
0.0712
0.2366
0.1738
0.2332
0.0831
0.0802
DISTS ↓
0.0724
0.0710
0.0666
0.0672
0.1460
0.0834
0.1166
0.0525
0.0429
HIDE
PSNR ↑
31.22
31.46
31.74
32.35
25.94
23.20
21.94
27.28
27.52
Table 1 : Quantitative comparison for single-image deblurring. Bold indicates the best result and underline indicates the second-best result among generative-based methods.
Figure 3 : Qualitative comparison for single-image deblurring on GoPro, HIDE, and RealBlur-J (top to bottom). Best viewed zoomed in.
Table 5
Figure 4 : Qualitative comparison for blur-to-video on GoPro-7. Best viewed zoomed in.
Deblurring
Blur-to-video
Method
PSNR ↑
SSIM ↑
LPIPS ↓
DISTS ↓
PSNR p ↑
SSIM p ↑
LPIPS p ↓
FVD ↓
EPE ↓
B→S
27.72
0.8944
0.1595
0.1149
29.16
0.8894
0.0292
59.51
1.15
B→{S(τi)}i=0N−1
27.76
0.8945
0.1604
0.1166
29.26
0.8933
0.0278
58.76
1.13
B→B (ours)
28.45
0.9080
0.1455
0.1043
29.78
0.9034
0.0271
59.29
1.02
Table 4 : Advantage of interval prediction on GoPro. B→B denotes our proposed method.
Figure 5 : Loss-term ablation visualization on GoPro-7 (frame 6 of 7). Best viewed zoomed in.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Backbone
FLUX.2-klein-4B (transformer only fine-tuned, VAE frozen)
Trainable params
∼ 4B (full transformer)
Optimizer
Adam
Total iterations
100K
LR schedule
Linear warmup to 10−4 over 2K, then cosine-annealing to 10−5
Mixed precision
bf16
Gradient accumulation
2 steps
Appendix
Table 6: Training hyperparameters.
PSNR (dB)
Description
Mean
Median
Std.
10th pct.
90th pct.
D(zs,t) vs. Bs,t
(VAE round-trip only)
42.70
43.12
2.32
39.27
45.24
D(zpred) vs. D(zs,t)
(pure additivity)
39.82
39.38
4.49
34.36
45.94
D(zpred) vs. Bs,t
(combined)
37.40
37.26
2.80
33.83
41.03
Appendix
Table 7: Latent-space additivity validity. PSNR distributions over N=2,983 random (s,t,m) triples drawn from 11 GoPro test sequences ( 512×512 crops). The pure additivity error D(zpred) vs. D(zs,t) is comparable to the VAE round-trip floor, indicating that E is approximately linear over the blur-image manifold.
Method
PSNR ↑
SSIM ↑
LPIPS ↓
DISTS ↓
Jin et al. [ 15 ]
29.44
0.9237
0.1345
0.0982
MotionETR [ 48 ]
28.60
0.9415
0.0509
0.0495
Blur2Vid [ 37 ]
30.66
0.9528
0.0363
0.0380
Ours
33.22
0.9664
0.0291
0.0298
Appendix
Table 8: Center frame quality on the GoPro-7 split. 1,744 clips at 640×360 . Frame index 3 of the recovered 7-frame sequence is evaluated against the corresponding ground-truth sharp frame using full-image metrics.
Dataset
Metric
Lrecon
Lrecon+Ladd
Lrecon+Lsharp
Full
GoPro
PSNR ↑
28.45
28.56
29.24
29.43
SSIM ↑
0.9080
0.9096
0.9236
0.9253
LPIPS ↓
0.1455
0.1422
0.0787
0.0802
DISTS ↓
0.1043
0.1020
0.0420
0.0429
HIDE
PSNR ↑
26.73
26.79
27.49
27.52
SSIM ↑
0.8668
0.8687
0.8894
0.8891
Appendix
Table 9: Cross-dataset loss-term ablation on single-image deblurring metrics. Extension of Table. 5 .
Figure 6: Qualitative comparison for the supervision-target ablation on GoPro-7. Best viewed zoomed in.
Figure 7: Additional blur-to-video comparisons on GoPro, trained on odd frames only.
Figure 8: Qualitative comparison for the loss-term ablation on GoPro-7 (sequence view). Each row shows a recovered 7 -frame sequence under a different combination of training losses. Best viewed zoomed in.
Figure 9: Qualitative comparison for the loss-term ablation on GoPro (center-frame deblurring view). Each column shows the recovered center sharp frame under a different combination of training losses. Best viewed zoomed in.
Figure 10: Additional single-image deblurring comparisons. Two rows per dataset, GoPro, HIDE, and RealBlur-J, RealBlur-R (top to bottom).
Figure 11: Additional blur-to-video comparisons on GoPro-7.
Asset
Type
Citation
License
GoPro_Large_All
Dataset
[ 28 ]
Research use (per authors)
HIDE
Dataset
[ 35 ]
Research use (per authors)
RealBlur-J / RealBlur-R
Dataset
[ 32 ]
CC BY-NC-SA 4.0
FLUX.2-klein
Pretrained model
[ 3 ]
FLUX.2 Community License (non-commercial)
Stable Diffusion (referenced)
Pretrained model
[ 33 ]
CreativeML Open RAIL-M
RAFT
Optical flow
[ 38 ]
BSD 3-Clause
Appendix
Table 10: Existing assets used in this work. Datasets, pretrained models, and software with their licenses.