cs.LGSep 30, 2026

Learning Where to Steer: Noise-Space Geometry for Efficient Offline Multi-Objective Optimization with Generative Models

Authors: Yuan Lu, Esha Singh, Yi-An Ma, Yusu Wang

Organizations: Halıcıoğlu Data Science Institute, UC San Diego · Department of Computer Science & Engineering, UC San Diego

Abstract

Offline multi-objective optimization (MOO) seeks solutions with better objective trade-offs using only a fixed dataset, without querying the objectives. Diffusion models trained on such data have emerged as a promising approach, but their samples are not inherently better than the data and must be steered toward the Pareto front. Existing methods guide or condition every sampling step. We instead act on the initial noise and leave the sampling process unchanged. Across Off-MOO-Bench, we observe that the objectives, as functions of the noise, are sensitive to only a few directions. We estimate these directions once per task via a Recursive Feature Machine using function values alone, and a small cache serves every trade-off, so each candidate costs one noise displacement and one ODE solve. We prove that this displacement increases the learned scalarized objective in expectation, and that sweeping trade-offs recovers the flow's attainable front up to proxy and steering errors. With additional guidance, for which we introduce novel data-adaptive and Pareto-aware operators, our method attains the best average hypervolume rank among generative methods on 47 tasks, at comparable or lower sampling cost. Steering alone outranks the best prior generative method at a fraction of its sampling cost.

Figures & tables

Appendix figures & tables39 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Jun 3, 2026cs.LG

ParetoPilot: Zero-Surrogate Offline Multi-Objective Optimization via Infer-Perturb-Guide Diffusion

Offline multi-objective optimization (Offline MOO) seeks Pareto-optimal designs from static datasets without additional environment interactions. Existing generative methods typically guide sampling with external surrogate or preference models, which adds training complexity and may provide unreliable guidance. We propose ParetoPilot, a plug-and-play method that guides designs to Pareto front at inference time using a pre-trained conditional diffusion model without any surrogate. ParetoPilot introduces an Infer-Perturb-Guide (IPG) engine within the reverse diffusion process. IPG first infers the individual conditional target for each sample in the batch by aligning its conditional and unconditional predictions. It then perturbs these targets collectively across the batch, balancing convergence toward the Pareto front and diversity among samples. Finally, the engine guides the generative trajectory toward the Pareto front by injecting these perturbed targets via standard Classifier-Free Guidance (CFG). Experiments on 51 tasks demonstrate that ParetoPilot achieves the best overall ranking among 16 methods and competitive hypervolume improvement.
Sep 7, 2026cs.LG

ParetoTransport: Generative Optimization by Mass Transport Toward The Pareto Front

Offline multi-objective optimization requires not only moving the objective vectors of candidate designs toward the Pareto front, but also distributing them effectively along it. Generative methods have recently emerged as a natural approach because they learn a distribution over feasible designs while allowing generation to be steered toward promising designs. Existing methods, however, largely retain classical sample-wise guidance strategies, leaving the distribution-level modeling capability of generative methods underused. We propose ParetoTransport, a training-free guidance method for pre-trained flow-matching models that explicitly specifies and refines a population-level distribution in objective space. ParetoTransport guides a flow-matching sampler to iteratively transport the empirical offline distribution toward the Pareto front, with Wasserstein matching to intermediate proxy distributions. This directly controls distributional displacement and mass allocation along the front. We establish a convergence result and demonstrate state-of-the-art performance on standard offline MOO benchmarks, extending recent evaluations beyond hypervolume to generational distance, inverted generational distance, and Wasserstein distance.
May 12, 2026cs.LG

Gradient-Free Noise Optimization for Reward Alignment in Generative Models

Existing reward alignment methods for diffusion and flow models rely on multi-step stochastic trajectories, making them difficult to extend to deterministic generators. A natural alternative is noise-space optimization, but existing approaches require backpropagation through the generator and reward pipeline, limiting applicability to differentiable settings. To address this, here we present ZeNO (Zeroth-order Noise Optimization), a gradient-free framework that formulates noise optimization as a path-integral control problem, estimable from zeroth-order reward evaluations alone. When instantiated with an Ornstein--Uhlenbeck reference process, the update connects to Langevin dynamics implicitly targeting a reward-tilted distribution. ZeNO enables effective inference-time scaling and demonstrates strong performance across diverse generators and reward functions, including a protein structure generation task where backpropagation is infeasible.