Fenchel Tilting: Weighted Correction for Efficient Finetuning of Generative Models
Organizations: AXXX · Applied AI Institute · MBZUAI
Abstract
Adapting a pretrained generative model to an arbitrary preference expressed as a utility function underlies reward alignment, guided design, and constraint satisfaction, enabling diverse applications. Existing fine-tuning methods trade off generality against computational cost: they either restrict the family class of supported preferences to keep optimization simple or preserve generality at the expense of efficiency. We introduce Fenchel Tilt Flow Control (FTFC), which decouples utility optimization from generative-model fitting. FTFC first optimizes for a target distribution by jointly fitting an effective reward and density-ratio weights on pretrained samples. Method combines the utility's variational structure with Fenchel duality, supporting general -divergence penalties that determine how rewards are transformed into an distribution-correction weights. These weights are then frozen and used to modify a diffusion or flow model in a single stage of importance-weighted denoising or flow matching, without differentiating through sampling trajectories. We establish exact duality for concave utilities under suitable conditions and show that weighted fitting reproduces the optimal target distribution for a given utility. Across image and molecule generation benchmarks, FTFC improves over baselines on diverse preference functions, while also being up to more efficient. roposed method enables adaptation beyond expected-reward maximization without complex optimization, while preserving robustness for more general class of the utility functions compared to baselines.
Figures & tables
| Method | Train days | Mean | SQ .998 | Upper | Valid. (%) | SA | Topology |
|---|---|---|---|---|---|---|---|
| Pretrained | |||||||
| AM | |||||||
| FDC | |||||||
| TFFT | |||||||
| FTFC | \bm{29.1}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.1} | \bm{39.4}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.3} | \bm{35.5}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.1} | 96.1\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 1.8} | \bm{4.4}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.1} | \bm{0.8}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.0} |
| ImageReward | Alignment | Diversity | Train | |||
|---|---|---|---|---|---|---|
| Method | Mean | L-CVaR .2 | CLIP | HPSv2.1 | DreamSim var. | days |
| Pretrained | 0.3\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.1} | -1.0\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.0} | 0.28\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.00} | 0.26\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.00} | 0.33\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.01} | |
| EXP-FT | 0.7\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.0} | -0.5\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.0} | 0.28\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.00} | \bm{0.28}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.00} | 0.30\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.01} | |
| FDC | 0.7\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.1} | -0.5\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.0} | 0.28\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.00} | 0.27\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.00} | 0.30\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.01} | |
| L-TFFT | 0.8\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.1} | -0.5\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.0} | 0.28\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.00} | 0.27\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.00} | 0.30\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.01} | |
| FTFC | \bm{0.9}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.1} | \bm{-0.4}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.0} | \bm{0.28}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.00} | 0.27\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.00} | \bm{0.35}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.01} | |
| Method | Mean | R-CVaR 0.9 | Valid. (%) | SA | Topology |
|---|---|---|---|---|---|
| Pretrained | 65.9\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.1} | 94.9\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.2} | 99.8\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.0} | 7.5\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.01} | 0.95\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.0} |
| EXP-FT | 67.9\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.8} | 94.7\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 1.5} | 99.8\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.2} | 7.3\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.03} | 0.93\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.0} |
| FDC | 69.2\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 2.4} | \bm{97.9}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 1.9} | 99.6\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.2} | 7.2\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.01} | 0.93\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.0} |
| R-TFFT | 70.6\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.8} | 90.8\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.9} | 99.8\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.2} | 7.3\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.03} | 0.93\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.0} |
| FTFC | \bm{75.3}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.2} | 95.0\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.3} | \bm{99.9}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.1} | \bm{7.1}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.01} | \bm{0.97}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.0} |
| penalty | Mean reward | Upper-tail CVaR | Topology |
|---|---|---|---|
| KL | \bm{75.3}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.2} | 95.0\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.1} | \bm{0.97}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.0} |
| Half-Pearson | 71.4\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.4} | 94.9\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.1} | 0.95\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.0} |
| Reverse KL | 68.0\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.4} | 93.8\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.1} | 0.95\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.0} |
| Cressie-Read-3 | 73.2\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.7} | \bm{96.3}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.1} | 0.96\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.0} |
| Squared Hellinger | 69.0\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.2} | 95.1\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.3} | 0.95\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.0} |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Utility | Representation | Reward / calibration |
| Expected reward | ; no features | |
| Moment matching | : chosen descriptors | |
| Coverage entropy | : region memberships | |
| D-optimal design , | ||
| Constraint barrier † | ||
| Lower-tail CVaR | Maximize over observed reward knots. |
| Setting | AM | FDC |
|---|---|---|
| Objective | Mean reward, | Upper SQ, , |
| Optimizer | Adam | Adam |
| Adam moments / epsilon | / | / |
| Learning rate | ||
| Optimizer updates | ||
| Outer iterations |
| Setting | Value |
|---|---|
| Runs per method | |
| Reward / KL coefficient | , |
| Target right-tail level | |
| Optimizer | AdamW; learning rate ; weight decay |
| Updates | EXP-FT / R-TFFT: ; FDC: |
| Effective batch / microbatch | / at most molecules |
| Method | XTP (%) | Reward given XTP | Strain given XTP |
|---|---|---|---|
| Pretrained | 98.52\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.20} | \bm{0.97}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.00} | \bm{0.18}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.00} |
| EXP-FT (AM) | 97.68\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.12} | 0.95\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.00} | 0.27\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.02} |
| FDC ( ) | 97.68\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 1.08} | 0.95\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.01} | 0.27\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.04} |
| R-TFFT | 98.23\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.26} | 0.94\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.02} | 0.32\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.11} |
| FTFC (ours) | \bm{98.67}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.08} | 0.97\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.00} | 0.19\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.00} |
| Method | Updates | Fitting + setup | Threshold refresh | Total training |
|---|---|---|---|---|
| Pretrained | ||||
| EXP-FT (AM) | 129.40\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.89} | 129.40\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.89} | ||
| FDC ( ) | 388.11\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 2.07} | 414.65\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 11.03} | 802.75\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 13.02} | |
| R-TFFT | 125.83\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.36} | 125.83\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.36} | ||
| FTFC (ours) | \bm{2.36}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.08} | \bm{2.36}\,{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 0.08} |
| Method | Epochs | Optimizer updates | Raw hinge threshold |
|---|---|---|---|
| Pretrained ( rombach2022latent ) | — | ||
| EXP-FT ( domingo2025adjoint ) | Identity reward | ||
| L-TFFT ( wang2026efficient ) | |||
| FDC, stage ( de2025flow ) | |||
| FDC, stage ( de2025flow ) |
| Setting | Value |
|---|---|
| Optimizer | AdamW; learning rate |
| Adam moments, epsilon | , |
| Weight decay | |
| Learning-rate schedule | Linear warmup for updates, then constant |
| Minibatch / accumulation | records / minibatches |
| Replay buffer / passes | trajectories / passes |