Coarse-guided visual generation, which synthesizes fine visual samples from degraded or low-fidelity coarse references, is essential for various real-world applications. While training-based approaches are effective, they are inherently limited by high training costs and restricted generalization due to paired data collection. Accordingly, recent training-free works propose to leverage pretrained diffusion models and incorporate guidance during the sampling process. However, these training-free methods either require knowing the forward (fine-to-coarse) transformation operator, e.g., bicubic downsampling, or are difficult to balance between guidance and synthetic quality. To address these challenges, we propose a novel guided method by using the h-transform, a tool that can constrain stochastic processes (e.g., sampling process) under desired conditions. Specifically, we modify the transition probability at each sampling timestep by adding to the original differential equation with a drift function h, which approximately steers the generation toward the ideal fine sample. To address unavoidable approximation errors, we introduce an adaptive weight scheduler that combines a noise-level-aware initialization with a correction based on cross-timestep consistency, balancing guidance adherence and synthesis quality. Extensive experiments across diverse image and video generation tasks demonstrate its effectiveness and generalization.
Figures & tables
Figure 1: Existing and our solutions. (a) Training translation networks based on paired data, which is costly and non-generalizable to different types of coarse samples. (b) Solving inverse problems based on a known forward operator, which is not robust. (c) Adding noise to the coarse sample and denoising it, which is difficult to balance guidance and quality. (d) Our method leverages the h -transform to achieve training-free, operator-free, and stable coarse-guided generation.
Figure 2: Overview of Weighted h -Transform Sampling . (a) If we have {\color[rgb]{0.7539,0,0}\bm{h}_{\bm{x}_{0}=\bm{y}}} , the generation result will be the ideal sample. (b) We leverage {\color[rgb]{0,0.6914,0.3125}\bm{h}_{\bm{x}_{0}=\widetilde{\bm{y}}}} to approximate the intractable {\color[rgb]{0.7539,0,0}\bm{h}_{\bm{x}_{0}=\bm{y}}} and derive that the error depends on the noise level and endpoint gap. (c) To mitigate the error influence, we adjust the approximation weight and finally generate a high-quality refined sample.
Method
Known Operator
SR
Inpaint
GD
MD
FID ↓
LPIPS ↓
FID ↓
LPIPS ↓
FID ↓
LPIPS ↓
FID ↓
LPIPS ↓
ADMM-TV
✓
110.6
0.428
181.5
0.463
186.7
0.507
152.3
0.508
ILVR ( Choi et al., 2021 ; Song et al., 2020 )
✓
96.72
0.563
76.54
0.612
109.0
0.403
292.2
0.657
PnP-ADMM ( Chan et al., 2016 )
✓
66.52
0.353
123.6
0.692
90.42
0.441
89.08
0.405
MCG ( Chung et al., 2022b )
✓
87.64
0.520
29.26
0.286
101.2
0.340
310.5
0.702
DDRM ( Kawar et al., 2022 )
✓
62.15
0.294
69.71
0.587
74.92
0.332
-
-
Table 1: Quantitative results of coarse-image guided generation on the FFHQ 256x256 validation dataset. Bold : best, underline : second best.
Method
MSE ↓
LPIPS ↓
FVD ↓
DINOv2 ↓
CLIP Cons. ↑
Optical Flow ↓
GT
-
-
-
-
0.974
-
Coarse Video
11.46
0.276
15.55
0.220
0.973
41.3
GWTF( γ=0.5 )
26.08
0.360
15.31
0.149
0.975
118.5
GWTF( γ=0.7 )
36.45
0.457
21.25
0.173
0.984
145.2
TTM( tw=4,ts=8 )
23.50
0.382
15.69
0.147
0.980
157.2
TTM( tw=4,ts=9 )
23.15
0.380
15.59
0.147
0.980
158.8
Table 2: Quantitative results of camera-controlled video generation on DL3DV.
Figure 3: Qualitative comparisons on the subset of DL3DV-10K. Our method shows better appearance alignment to the ground truth (see highlighted blue boxes).
λ
SR
Inpaint
GD
MD
0.0
0.586
0.586
0.586
0.586
0.5
0.456
0.989
0.445
0.491
1.0
0.472
1.008
0.483
0.546
2.0
0.466
1.009
0.486
0.552
sin(2Tπt)
0.224
0.448
0.293
0.344
Tt
0.223
0.406
0.274
0.344
Table 3: The ablation of weight schedule λ .
Method
Source Consistency
Semantic Alignment
Distance ↓
PSNR ↑
LPIPS ↓
MSE ↓
SSIM ↑
CLIP Entire ↑
CLIP Edited ↑
ODE-Inv
0.074
17.57
0.240
0.024
0.691
24.57
21.73
SDEdit
0.036
22.57
0.119
0.008
0.747
24.56
21.95
iRFDS
0.069
18.81
0.191
0.021
0.738
25.12
21.95
FlowEdit
0.036
23.02
0.082
0.007
0.842
25.98
22.81
FlowAlign
0.028
25.50
0.053
0.004
0.879
25.28
22.00
Table 4: Quantitative comparisons with five baselines for image editing.
Figure 4: Generalize to video motion transfer.
Figure A5: Our method is compatible with both score-based CogVideoX and Flow-based Wan2.2.
Figure A6: Super-resolution results by using different α . A large α leads to a small weight, and vice versa (weight scheduler λσ=σtα ).
Figure A7: Additional qualitative results on image restoration tasks.
Figure A8: Additional qualitative comparisons with SDEdit on image restoration tasks.
Figure A9: Additional camera-controlled video generation qualitative comparisons.
Figure A10: Additional camera-controlled video generation results using Wan2.2.
Figure A11: Text-based image editing comparisons with three baselines on PIE-Bench