cs.LGSep 26, 2026

Timestep Weighting: A Hidden Key to Effective ELBO-Based Flow-Matching RL

Authors: Qinwei Ma, Jingzhe Shi, Simin Fan, Ling Li, Alex Lamb

Organizations: College of AI, Tsinghua University · Department of Electrical and Computer Engineering, Princeton University · EPFL · Computer Science Department, Tsinghua University

Abstract

ELBO-based reinforcement learning offers a sampler-agnostic approach to fine-tuning flow matching models with reward feedback. Timestep weighting in ELBO-based RL has large impact on performance, and it also provides a unified view (as we show in this work) to understand prediction losses heuristically chosen in prior work, yet it remains under-researched and is often chosen to inherit pretrain configs. We investigate impacts and dynamics of timestep weighting in ELBO-based RL. We show that effective weighting depends on both the reward landscape and stage of learning. (1) Through experiments on controlled CIFAR image generation, complemented by robotics, we investigate how weighting impacts reward-driven updates across noise levels. (2) Through gradient analysis, we reveal distinct patterns of cross-noise coordination across tasks and their evolution during training. These findings motivate the hypothesis that useful weighting depends on the gap between the policy's current behavior and the behavior favored by the reward. (3) Guided by this analysis, we study simple static weighting, budgeted profile selection, and dynamic schedules that improve performance beyond conventional target choices. Our results establish timestep weighting as an important design choice for flow-matching RL and motivate further research into methods that choose and adapt it throughout learning.

Figures & tables

Appendix figures & tables41 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Precise: SDE-Consistent Stochastic Sampling for RL Post-Training of Flow-Matching Models

    May 22, 2026Jade Zou, Tao Huang, Weijie Kong +7Accelerated SamplingReward Gradients

  2. Reinforce Adjoint Matching: Scaling RL Post-Training of Diffusion and Flow-Matching Models

    May 11, 2026Andreas Bergmeister, Stefanie Jegelka, Nikolas Nüsken +2Generative Flow NetworksDiffusion Models

  3. The Reward Was in Your Data All Along: Correcting Flow Matching with Discriminator-Guided RL

    Jun 17, 2026Nicolas Beltran-Velez, Felix Friedrich, Zhang Xiaofeng +4Contrastive Flow MatchingDiffusion Alignment