Recent advances in distillation and flow-map models have enabled deterministic one- or few-step generation for high-quality data, facilitating a new branch of reward alignment approaches that directly optimize the initial noise from a Gaussian distribution. However, most existing initial-noise optimization methods rely on first-order gradient information, which is either inapplicable or suffers from instability and inefficiency in black-box reward scenarios. Here, we introduce ZeNOVA, a stable and efficient initial noise alignment method in a gradient-free manner. Specifically, we address existing algorithms' major challenge in black-box scenarios through annealed soft-value guidance, manifold-constrained hyperspherical Langevin dynamics, and Metropolis-Hastings jumping. Extensive experiments on image and video generative models show that ZeNOVA outperforms all evaluated zeroth-order baselines by optimizing the initial noise toward higher rewards substantially more stably while exploiting the geometry of the Gaussian prior, demonstrating its practical applicability to various black-box reward alignment.
Figures & tables
Figure 1: Overview of ZeNOVA. Naive initial-noise optimization suffers from limited exploration and unstable updates under inaccurate zeroth-order guidance from raw reward signal. ZeNOVA combines annealed soft-value guidance, hyperspherical updates, and Metropolis-Hastings jumping to enable stable local refinement and global exploration while respecting the Gaussian prior geometry.
Figure 2
Figure 4: Qualitative image results from SDXL-Turbo and DMD2.
ImageReward
PickScore
GenEval
Counting
OCR
Method
Reward ↑
Aes. ↑
Val. ↑
Reward ↑
Aes. ↑
Val. ↑
Reward ↑
Aes. ↑
Reward ↓
Aes. ↑
Reward ↑
Aes. ↑
SDXL-Turbo
base model
1.0505
6.0282
22.8809
21.5378
5.4020
0.5321
0.5375
5.3853
28.375
5.6965
0.1368
5.4256
Best-of- K
1.6932
6.1148
23.2829
23.1925
5.5747
0.8949
0.7232
5.3717
2.9
5.7096
0.4784
5.3303
ReNO
1.7378
6.0395
22.9585
25.1031
5.5721
0.7755
-
-
-
-
-
-
ORIGEN
1.8198
5.1640
21.1789
24.2321
5.5748
0.9286
-
-
-
-
-
-
Table 1: Initial noise reward alignment performance on text-to-image models. Bold : best performance, “-”: Not available.
Figure 5: Qualitative video results from rCM with ZeNOVA, optimizing VideoAlign scores.
VideoAlign
Method
VQ
MQ
TA
All
base model
3.4854
1.3684
2.1731
6.1997
ZeNOVA
6.4464
3.3124
3.9088
10.8146
Table 2: Quantitative performance on text-to-video models. Bold : Best performance.
Figure 7
Variants
ImageReward
PickScore
OCR
ZeNOVA
1.8078
23.4564
0.5114
- MH jumping
1.7262
23.1867
0.4412
- Value Annealing
1.6984
23.2960
0.3756
- Hyperspherical update
1.3968
22.2734
0.3823
Table 3: Ablation study of the components of ZeNOVA.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Images generated by SDXL-Turbo and DMD2 with ZeNOVA, using ImageReward as the reward signal.
Figure 9: Images generated by SDXL-Turbo and DMD2 with ZeNOVA, using PickScore as the reward signal.
Figure 10: Images generated by SDXL-Turbo and DMD2 with ZeNOVA, using GenEval as the reward signal.
Figure 11: Images generated by SDXL-Turbo and DMD2 with ZeNOVA, using the counting reward as the reward signal.
Figure 12: Images generated by SDXL-Turbo and DMD2 with ZeNOVA, using the OCR reward as the reward signal.
Figure 13: Vidoes generated by rCM with ZeNOVA, using a sum of three VideoAlign rewards as the reward signal.
Existing reward alignment methods for diffusion and flow models rely on multi-step stochastic trajectories, making them difficult to extend to deterministic generators. A natural alternative is noise-space optimization, but existing approaches require backpropagation through the generator and reward pipeline, limiting applicability to differentiable settings. To address this, here we present ZeNO (Zeroth-order Noise Optimization), a gradient-free framework that formulates noise optimization as a path-integral control problem, estimable from zeroth-order reward evaluations alone. When instantiated with an Ornstein--Uhlenbeck reference process, the update connects to Langevin dynamics implicitly targeting a reward-tilted distribution. ZeNO enables effective inference-time scaling and demonstrates strong performance across diverse generators and reward functions, including a protein structure generation task where backpropagation is infeasible.
Optimizing the noise samples of diffusion and flow models is an increasingly popular approach to align these models to target rewards at inference time. However, we observe that these approaches are usually restricted to differentiable or cheap reward models, the formulation of the underlying pretrained generative model, or are memory/compute inefficient. We instead propose a simple trust-region based search algorithm (TRS) which treats the pre-trained generative and reward models as a black-box and only optimizes the source noise. Our approach achieves a good balance between global exploration and local exploitation, and is versatile and easily adaptable to various generative settings and reward models with minimal hyperparameter tuning. We evaluate TRS across text-to-image, molecule and protein design tasks, and obtain significantly improved output samples over the base generative models and other inference-time alignment approaches which optimize the source noise sample, or even the entire reverse-time sampling noise trajectories in the case of diffusion models. Our source code is publicly available.
Niklas Schweiger, Daniel Cremers, Karnik Ram
Technical University of Munich, Germany · Munich Center for Machine Learning, Germany
Diffusion and flow-matching models are typically trained by corrupting data through independently sampled Gaussian noise. While simple and scalable, this forward process induces arbitrary data-noise couplings, forcing the network to learn high-curvature transports between unrelated endpoints. Existing optimal-transport methods reduce this burden by reassigning fixed noise samples to data, but the source noise distribution itself remains passive. To address this, we introduce Contrastive Noise Alignment (CNA), a training-time method that creates dynamic, contrastive couplings by optimizing the noise representations directly. By modeling the noise batch as an interacting particle system, CNA employs a cross-modal InfoNCE objective to align noise particles with their paired data targets. To prevent spatial collapse, this alignment is regularized using an angular entropy term and a radial norm penalty. We show theoretically that this equilibrium asymptotically preserves Gaussian structures, maintaining tractability during inference. Empirically, CNA improves the alignment between noise and data, reduces flow curvature, and provides better generation quality with fewer required sampling steps. For few-step, pixel-space generation (2-4 NFEs), CNA reduces FID by over 50% compared to standard rectified flow, and by at least 24% against Optimal Transport baselines.