Few-step generative models can generate high-fidelity samples within a few function evaluations. Despite this efficiency, generated samples may not exhibit desirable properties. When these properties are difficult to encode as an explicit reward function, direct preference optimization (DPO) can align generative models using pairwise preference feedback without training a separate reward model. However, extending DPO to few-step generative models is challenging because few-step generative models are generally implicit, making the likelihood evaluation required by DPO intractable. To address this challenge, we introduce Few-step DPO (FestDPO), an extension of DPO for few-step generative models that leverages nonparametric likelihood estimation from empirical samples. By exploiting the fast sampling capabilities of few-step generative models, our approach makes sample-based approximation of DPO loss computationally feasible. Furthermore, the sample-based formulation makes FestDPO agnostic to the model family and sampling procedure. Our toy experiment demonstrates that FestDPO matches the reward-tilted target distribution across four few-step generators. For real-world tasks, we evaluate FestDPO in two domains: text-to-image generation and protein backbone generation. In text-to-image generation, FestDPO outperforms preference optimization baselines in both win rates against the base models and human evaluation scores. In protein backbone generation, it achieves a higher β-sheet fraction and better structural designability than the baselines.
Figures & tables
Figure 1: Qualitative comparison of PSO ( Miao et al., 2025 ) and FestDPO on three text prompts: “a gorgeous queen with cat-like eyes” (Left), “ three chairs” (Middle), and “3D digital illustration, burger with wheels speeding on the race track, supercharged, detailed, hyperrealistic, 4K” (Right).
Figure 2: Overview of FestDPO. For a prompt x and a preference pair (yw≻yl) , we draw N samples from both the trainable policy πθ and the frozen reference policy πref . We then construct kernel density estimates π^θ and π^ref and evaluate them at yw and yl . These estimates yield a tractable surrogate for the DPO loss ( Equation 8 ), which we minimize to align πθ with the preference data.
Figure 3: FestDPO aligns diverse few-step generators. (a) The reference distribution pref(x) and (b) the reward-tilted target distribution p⋆(x)∝pref(x)exp(r⋆(x)) . (c)–(f) Distributions of Drifting, IMM, MeanFlow, and sCM fine-tuned with FestDPO. Across all four generators, the fine-tuned distributions closely match p⋆(x) , with KL divergences to the target of approximately 3×10−3 .
Figure 4: Win rate comparison against (a) SDXL-Turbo and (b) SDXL-DMD2 on Pick-a-Pic v2, Parti-Prompts, and HPSv2 evaluation datasets. Each bar represents the mean win rate (%) of each method for PickScore and Aesthetic score, and error bars denote the standard deviation.
Figure 5: PickScore of DrPO, PSO, and FestDPO during fine-tuning of SDXL-DMD2 on Pick-a-Pic v2, Parti-Prompts, and HPSv2. Lines and shaded regions denote the mean and standard deviation.
Table 6
Method
Align. ↑
Aesth. ↑
SDXL-Turbo
3.683
3.149
PSO
3.767
3.286
DrPO
3.696
3.158
FestDPO (ours)
4.199
3.860
Table 3: Human evaluation results (mean rating). Best results are in bold .
Table 8
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: FestDPO aligns diverse few-step generators with the target distribution. Left of the solid line: the reference distribution pref(x) , a mixture of Gaussians, and the reward-tilted target distribution p⋆(x)∝pref(x)exp(r⋆(x)) . Right of the solid line: distributions of (a) Drifting model, (b) IMM, (c) MeanFlow, and (d) sCM) before (top) and after (bottom) fine-tuning with FestDPO. Dashed and solid curves denote pref(x) and p⋆(x) , respectively. We report the KL divergence (KL) between the fine-tuned and target distributions.
Drifting
IMM
MeanFlow
sCM
Training steps
300k
100k
300k
100k
Optimizer
AdamW
RAdam
Adam
AdamW
Learning rate
2×10−4
1×10−4
2×10−4
3×10−5
(β1,β2)
(0.9, 0.95)
(0.9, 0.999)
(0.9, 0.999)
(0.9, 0.99)
Weight decay
0.01
0
0
0.01
Batch size
1024
4096
1024
16384
Appendix
Table 6: Pre-training hyperparameters.
Drifting
IMM
MeanFlow
SCM
Fine-tuning steps
4,000
4,000
8,000
4,000
Learning rate
1×10−4
1×10−4
3×10−5
1×10−4
β
1.0
1.0
1.0
1.0
KDE bandwidths
0.02 / 0.06 / 0.2
0.02 / 0.06 / 0.2
0.01 / 0.03 / 0.1
0.02 / 0.06 / 0.2
# Samples ( M=N )
2048
2048
8192
2048
Pairs per step
512
512
512
512
Appendix
Table 7: Fine-tuning hyperparameters of FestDPO.
SDXL-Turbo
SDXL-DMD2
KDE bandwidth h
0.15, 0.2
0.2, 0.3
β
0.5
# of KDE samples N
12
LoRA rank / scale
16 / 1
Optimizer
AdamW ( β1=0.9,β2=0.999 )
Weight decay
0.01
Appendix
Table 8: Hyperparameters for FestDPO.
Hyperparameter
SS-match
scRMSD
KDE bandwidth σ
0.5
0.05
DPO coefficient β
2000
20
Base model
RMF-S
KDE samples per update
64 trainable / 64 reference
Optimizer
AdamW
AdamW (β1,β2)
(0.9,0.999)
Appendix
Table 9: Hyperparameters and candidate values for FestDPO on protein backbone generation.
Pick-a-Pic v2
Parti-Prompts
HPSv2
Method
PickScore ↑
Aesthetic ↑
PickScore ↑
Aesthetic ↑
PickScore ↑
Aesthetic ↑
SDXL-Turbo
22.37 ± 0.06
6.03 ± 0.03
22.78 ± 0.04
5.69 ± 0.02
22.83 ± 0.05
6.09 ± 0.02
DrPO
22.38 ± 0.06
6.03 ± 0.03
22.78 ± 0.04
5.69 ± 0.02
22.85 ± 0.05
6.09 ± 0.02
PSO
22.46 ± 0.06
6.06 ± 0.02
22.86 ± 0.04
5.72 ± 0.02
23.22 ± 0.05
6.05 ± 0.02
FestDPO
22.60 ± 0.07
6.07 ± 0.03
22.92 ± 0.05
5.72 ± 0.02
23.18 ± 0.05
6.14 ± 0.02
Appendix
Table 10: Main results with SDXL-Turbo on Pick-a-Pic v2, Parti-Prompts, HPSv2.
Pick-a-Pic v2
Parti-Prompts
HPSv2
Method
PickScore ↑
Aesthetic ↑
PickScore ↑
Aesthetic ↑
PickScore ↑
Aesthetic ↑
DrPO
51.65 ± 1.21
49.40 ± 1.27
50.20 ± 1.00
53.65 ± 1.02
54.50 ± 1.13
51.30 ± 1.09
PSO
65.60 ± 1.29
56.20 ± 1.35
64.70 ± 1.09
57.05 ± 1.08
77.70 ± 1.03
44.35 ± 1.33
FestDPO
73.55 ± 1.24
59.45 ± 1.27
67.65 ± 1.08
63.05 ± 1.04
82.25 ± 0.96
61.00 ± 1.18
Appendix
Table 11: Win-rate comparison with SDXL-Turbo on Pick-a-Pic v2, Parti-Prompts, HPSv2.
Pick-a-Pic v2
Parti-Prompts
HPSv2
Method
PickScore ↑
Aesthetic ↑
PickScore ↑
Aesthetic ↑
PickScore ↑
Aesthetic ↑
SDXL-DMD2
21.06 ± 1.46
5.37 ± 0.68
21.65 ± 1.17
5.25 ± 0.57
21.45 ± 1.28
5.58 ± 0.67
DrPO
21.17 ± 1.47
5.43 ± 0.70
21.73 ± 1.19
5.31 ± 0.57
21.61 ± 1.31
5.63 ± 0.68
PSO
21.58 ± 1.54
5.66 ± 0.70
22.06 ± 1.35
5.51 ± 0.64
22.05 ± 1.38
5.90 ± 0.62
FestDPO
21.98 ± 1.50
5.87 ± 0.57
22.43 ± 1.16
5.65 ± 0.52
22.42 ± 1.32
5.99 ± 0.59
Appendix
Table 12: Main results with SDXL-DMD2 on Pick-a-Pic v2, Parti-Prompts, HPSv2.
Pick-a-Pic v2
Parti-Prompts
HPSv2
Method
PickScore ↑
Aesthetic ↑
PickScore ↑
Aesthetic ↑
PickScore ↑
Aesthetic ↑
DrPO
67.00 ± 1.67
61.55 ± 1.98
63.15 ± 2.29
63.00 ± 1.99
72.00 ± 1.21
63.35 ± 1.19
PSO
77.00 ± 0.42
76.15 ± 0.43
73.85 ± 0.44
73.70 ± 0.44
78.90 ± 0.41
78.60 ± 0.41
FestDPO
88.30 ± 0.72
87.25 ± 0.75
86.55 ± 0.76
84.60 ± 0.81
90.90 ± 0.64
85.05 ± 0.80
Appendix
Table 13: Win-rate comparison with SDXL-DMD2 on Pick-a-Pic v2, Parti-Prompts, HPSv2.
Pick-a-Pic v2
Parti-Prompts
HPSv2
# of KDE samples
PickScore ↑
Aesthetic ↑
PickScore ↑
Aesthetic ↑
PickScore ↑
Aesthetic ↑
SDXL-Turbo
22.37 ± 1.46
6.03 ± 0.61
22.78 ± 1.20
5.69 ± 0.57
22.83 ± 1.29
6.09 ± 0.69
N=8
22.53 ± 1.51
6.06 ± 0.62
22.91 ± 1.24
5.74 ± 0.58
23.06 ± 1.32
6.13 ± 0.68
N=12
22.60 ± 1.52
6.07 ± 0.62
22.92 ± 1.25
5.73 ± 0.59
23.13 ± 1.35
6.11 ± 0.66
N=24
22.61 ± 1.50
6.02 ± 0.60
22.94 ± 1.24
5.72 ± 0.57
23.14 ± 1.37
6.07 ± 0.63
Appendix
Table 14: Sensitivity analysis on the number of KDE samples with SDXL-Turbo.
Pick-a-Pic v2
Parti-Prompts
HPSv2
Feature encoder
PickScore ↑
Aesthetic ↑
PickScore ↑
Aesthetic ↑
PickScore ↑
Aesthetic ↑
SDXL-Turbo
22.37 ± 1.46
6.03 ± 0.61
22.78 ± 1.20
5.69 ± 0.57
22.83 ± 1.29
6.09 ± 0.69
Latent MAE
22.60 ± 1.52
6.07 ± 0.62
22.92 ± 1.25
5.73 ± 0.59
23.13 ± 1.35
6.11 ± 0.66
CLIP
22.66 ± 1.48
5.99 ± 0.57
22.91 ± 1.22
5.72 ± 0.56
23.17 ± 1.35
6.08 ± 0.64
DINOv2
22.54 ± 1.49
6.02 ± 0.61
22.92 ± 1.22
5.71 ± 0.59
23.07 ± 1.34
6.08 ± 0.65
Appendix
Table 15: Ablation on the feature encoder with SDXL-Turbo.
Pick-a-Pic v2
Parti-Prompts
HPSv2
NFE
PickScore ↑
Aesthetic ↑
PickScore ↑
Aesthetic ↑
PickScore ↑
Aesthetic ↑
1
22.60 ± 1.52
6.07 ± 0.62
22.92 ± 1.25
5.73 ± 0.59
23.13 ± 1.35
6.11 ± 0.66
2
22.63 ± 1.48
5.98 ± 0.59
23.07 ± 1.22
5.67 ± 0.54
23.14 ± 1.33
6.06 ± 0.67
Appendix
Table 16: Performance of FestDPO with SDXL-Turbo under different NFE.
Figure 8: Qualitative comparison of FestDPO with PSO and DrPO on Pick-a-Pic V2.
Figure 9: Qualitative comparison of FestDPO with PSO and DrPO on Parti-Prompts.
Figure 10: Qualitative comparison of FestDPO with PSO and DrPO on HPSv2.
Figure 11: Screenshot of the annotation interface used for human evaluation. For each text prompt, four images generated by different methods are displayed side by side in randomized order. Annotators rate each image on a 1–5 scale for prompt alignment and aesthetics.
Overall
Pick-a-Pic v2
Parti-Prompts
HPSv2
Method
Align. ↑
Aes. ↑
Align. ↑
Aes. ↑
Align. ↑
Aes. ↑
Align. ↑
Aes. ↑
SDXL-Turbo
3.683
3.149
3.756
3.248
3.497
3.080
3.735
3.101
PSO
3.767
3.286
3.731
3.156
3.607
3.103
3.903
3.525
DrPO
3.696
3.158
3.759
3.234
3.520
3.057
3.751
3.151
FestDPO (ours)
4.199
3.860
4.329
3.936
4.137
3.793
4.118
3.832
Appendix
Table 17: Human evaluation results on Pick-a-Pic v2, Parti-Prompts, HPSv2.
Figure 12: Pairwise win rates (%) in the human evaluation for prompt alignment and aesthetics. Each entry indicates the percentage of comparisons in which the method on the y-axis is preferred over the method on the x-axis.
One-step text-to-image generators are attractive for deployment because they generate an image with a single forward pass, but preference finetuning them remains difficult: standard alignment methods often rely on policy likelihoods, denoising trajectories, differentiable reward gradients, or test-time optimization. We propose Drifting Preference Optimization (DrPO), an online preference-finetuning method for deterministic one-step generators. For each prompt, DrPO samples candidates from the current generator, ranks them with a target reward, and uses high- and low-scoring samples to synthesize a feature-space update direction. The update is a non-parametric dipole preference field plus a reference drift estimated from the frozen base generator, and is optimized through a detached feature-space regression target. The target reward is used only for ranking, so DrPO can train with large, black-box, or non-differentiable rewards while inference remains a single generator call. We evaluate DrPO on SD-Turbo and SDXL-Turbo with multiple target rewards and benchmarks, including HPSv3 and GenEval. DrPO improves alignment over reward-gradient-free one-step preference baselines and reduces HPSv3 training computation by 3.51× under the matched effective-batch setting by removing reward-model backpropagation. Initial offline experiments suggest that sample-based gradient synthesis can also be used beyond online reward ranking.
Zhou Jiang, Yandong Wen, Zhen Liu
Westlake University · The Chinese University of Hong Kong, Shenzhen
Achieving high-fidelity generation in extremely few sampling steps has long been a central goal of generative modeling. Existing approaches largely rely on distillation-based frameworks to compress the original multi-step denoising process into a few-step generator. However, such methods inherently constrain the student to imitate a stronger multi-step teacher, imposing the teacher as an upper bound on student performance. We argue that introducing \textbf{preference alignment awareness} enables the student to optimize toward reward-preferred generation quality, potentially surpassing the teacher instead of being restricted to rigid teacher imitation. To this end, we propose \textbf{Reward-Aware Trajectory Shaping (RATS)}, a lightweight framework for preference-aligned few-step generation. Specifically, teacher and student latent trajectories are aligned at key denoising stages through horizon matching, while a \textbf{reward-aware gate} is introduced to adaptively regulate teacher guidance based on their relative reward performance. Trajectory shaping is strengthened when the teacher achieves higher rewards, and relaxed when the student matches or surpasses the teacher, thereby enabling continued reward-driven improvement. By seamlessly integrating trajectory distillation, reward-aware gating, and preference alignment, RATS effectively transfers preference-relevant knowledge from high-step generators without incurring additional test-time computational overhead. Experimental results demonstrate that RATS substantially improves the efficiency--quality trade-off in few-step visual generation, significantly narrowing the gap between few-step students and stronger multi-step generators.
Rui Li, Bingyu Li, Yuanzhi Liang +3
University of Science and Technology of China HeFei, China · TeleAI ShangHai, China
Aligning a few-step generative model is challenging, since existing alignment frameworks typically rely on restrictive assumptions: a tractable likelihood, a specific ODE/SDE solver, or a particular model family. We introduce FAV, Few-step Generative Models Alignment via Sample-based Variational Inference, a general alignment framework that requires only sample access to the generator and the reference distribution. We cast alignment as sampling from a reward-tilted distribution anchored to a reference distribution. We leverage Stein Variational Gradient Descent as a sample-based variational inference scheme and amortize its particle updates into the generator parameters via fixed-point regression. We evaluate FAV on two domains: robotics manipulation and image generator alignment. On generative policy alignment for robotic manipulation, FAV outperforms prevailing policy extraction baselines across 56 offline and 30 offline-to-online RL tasks. For image generator alignment, FAV fine-tunes diverse few-step backbones, including GAN, drifting model, consistency models, and flow maps, scaling from ImageNet-256 to 10242 text-to-image synthesis. Code is available at https://github.com/Jaewoopudding/FAV.
Jaewoo Lee, Hyeongyu Kang, Dohyun Kim +9
1KAIST · 2MongooseAI · 3Mila – Quebec AI Institute +3