Timestep Weighting: A Hidden Key to Effective ELBO-Based Flow-Matching RL
Organizations: College of AI, Tsinghua University · Department of Electrical and Computer Engineering, Princeton University · EPFL · Computer Science Department, Tsinghua University
Abstract
ELBO-based reinforcement learning offers a sampler-agnostic approach to fine-tuning flow matching models with reward feedback. Timestep weighting in ELBO-based RL has large impact on performance, and it also provides a unified view (as we show in this work) to understand prediction losses heuristically chosen in prior work, yet it remains under-researched and is often chosen to inherit pretrain configs. We investigate impacts and dynamics of timestep weighting in ELBO-based RL. We show that effective weighting depends on both the reward landscape and stage of learning. (1) Through experiments on controlled CIFAR image generation, complemented by robotics, we investigate how weighting impacts reward-driven updates across noise levels. (2) Through gradient analysis, we reveal distinct patterns of cross-noise coordination across tasks and their evolution during training. These findings motivate the hypothesis that useful weighting depends on the gap between the policy's current behavior and the behavior favored by the reward. (3) Guided by this analysis, we study simple static weighting, budgeted profile selection, and dynamic schedules that improve performance beyond conventional target choices. Our results establish timestep weighting as an important design choice for flow-matching RL and motivate further research into methods that choose and adapt it throughout learning.
Figures & tables
| 1 | for do |
|---|---|
| 2 | . |
| 3 | Weighting: , , . |
| 4 | Collect , score , and compute stopped GRPO advantages . |
| 5 | For each , store pairs and old losses . |
| 6 | Recompute current losses on the same endpoints, times, and noises. |
| 7 | Compute and by Equation ( 6 ), then by Equation ( 7 ). |
Appendix figures & tables41 assets
Supplementary material from the paper’s appendix.
Appendix
| Objective | Profile placement | Residual and reduction |
|---|---|---|
| Calibrated CIFAR | Sample ; scalar inside exponent | Velocity error / ; aggregate, exponentiate, PPO clip |
| FPO | Explicit inside exponent | Noise error; aggregate, exponentiate, PPO clip |
| FPO++ | Explicit outside per-draw proxy | Velocity error / ; process each draw, then average |
| Reward | Profile–seed pairs | Peak difference explicit | W/T/L |
|---|---|---|---|
| Auto-A35 | 27 | 13/5/9 | |
| Edge | 27 | 23/4/0 |
| Reward | Max regret | |||||
|---|---|---|---|---|---|---|
| Auto-A35 | 0.0000 | |||||
| Mixed | 0.1268 | |||||
| Edge | 0.0000 | |||||
| CLIP | 0.0000 | |||||
| SimCLR | 0.0000 | |||||
| Digit7 | 0.0248 |
| Reward | DD (%) | Explicit | cohort | |||
|---|---|---|---|---|---|---|
| Automobile | 49.87 | .005621 | .000733 | .108042 | ||
| Auto-A35 | 48.70 | .004578 | .000638 | .018274 | .128433 | |
| CLIP | 36.85 | .010949 | .009414 | .035520 | .091558 | |
| SimCLR | 40.76 | .016113 | .015405 | .027671 | .052726 | |
| Digit7 | 46.35 | .008746 | .008322 | .002082 | — | |
| Edge | 6.25 | .000209 | .010453 | .011644 |
| Statistic | Value |
|---|---|
| Digit7 , updates 20–40 | |
| Digit7 , updates 20–80 | |
| Digit7 , updates 20–200 | |
| Explicit radius 4, : | |
| Leave-one-reward , explicit | |
| Leave-one-reward , cohort |
| Cohort | Radius | ||||
|---|---|---|---|---|---|
| Explicit | 2 | ||||
| Explicit | 4 | ||||
| Explicit | 8 | ||||
| cohort | 2 | ||||
| cohort | 4 | ||||
| cohort | 8 |
| Task | Checkpoint | Training seeds | Peak reward/return | |
|---|---|---|---|---|
| SimCLR | 200 | 20260724 / 20260724 | ||
| CLIP | 200 | 20260724 / 20260724 | ||
| Digit7 | 200 | 20260724 / 20260726 | ||
| Go2 | 300 | 9200, 9201 (both rows) |
| Training seed | Update | Training reward | Cross-bin mean | |
|---|---|---|---|---|
| -2.0000 | 20260725 | 100 | 0.3472 | 0.1448 |
| -1.0000 | 20260724 | 100 | 0.3368 | 0.1520 |
| 0.0000 | 20260724 | 100 | 0.3353 | 0.1787 |
| 1.0000 | 20260726 | 100 | 0.2886 | 0.0778 |
| Task | High–high | Low–low | High–low |
|---|---|---|---|
| Walker | |||
| Swimmer | |||
| Go2 |
| Method | Task | Checkpoints | Median within–cross gap |
|---|---|---|---|
| FPO | Acrobot | 24 | 0.0724 |
| FPO | Ball in Cup | 23 | 0.1050 |
| FPO | Cheetah | 22 | 0.0668 |
| FPO | Fish | 24 | 0.1055 |
| FPO | Swimmer | 24 | 0.0461 |
| FPO | Walker | 24 | 0.0453 |
| Quantity | All 80 | Omit two |
|---|---|---|
| 75/80 | 73/78 | |
| Mean off-diagonal | 80/80 | 78/78 |
| Median | 0.17480 | 0.15766 |
| Median mean off-diagonal | 0.09156 | 0.09335 |
| Median fraction of positive off-diagonal entries | 80.30% | 80.30% |
| States with negative off-diagonal tenth percentile | 68/80 | 66/78 |
| Reward | States | Median | Median mean |
|---|---|---|---|
| Auto-A35 | 20 | 0.09546 | 0.09064 |
| CLIP | 20 | 0.19128 | 0.09596 |
| SimCLR | 20 | 0.08243 | 0.08937 |
| Digit7 | 20 | 0.40711 | 0.09098 |
| Reward | Rule | Peak validation reward |
|---|---|---|
| Auto-A35 | Static direct | |
| Auto-A35 | Online coherence | |
| Mixed | Static direct | |
| Mixed | Online coherence | |
| Edge | Static direct | |
| Edge | Online coherence |
| Reward | Peak validation reward | Updates | |
|---|---|---|---|
| Auto-A35 | 200–280 | ||
| CLIP | 200–440 | ||
| Digit7 | 200–300 |
| Reward | Allocation | Observed peak | Step 800 |
|---|---|---|---|
| CLIP | Fixed | — | |
| CLIP | Fixed | — | |
| CLIP | Fixed | — | |
| CLIP | |||
| CLIP | Static softmax | — | |
| Digit7 | Fixed | — |
| Task | Native | Selected | Oracle | Gap | |
|---|---|---|---|---|---|
| Acrobot | 163.631157 | 185.187285 | 191.103217 | +21.556128 | 5.915933 |
| Spot | 313.813333 | 233.430000 | 313.813333 | -80.383333 | 80.383333 |
| Walker | 724.802788 | 742.842999 | 754.124350 | +18.040210 | 11.281352 |
| Task | First cut | Selected score | Exact draw range | |
|---|---|---|---|---|
| Acrobot | 10% | 182.14 | [176.86, 191.10] | 4.1803 |
| Acrobot | 20% | 185.19 | [176.86, 191.10] | 5.3607 |
| Acrobot | 40% | 188.88 | [176.86, 191.10] | 7.7582 |
| Acrobot | Full search | 187.30 | [176.86, 191.10] | 12.0000 |
| Spot | 10% | 197.73 | [197.73, 197.73] | 3.8060 |
| Spot | 20% | 233.43 | [233.43, 233.43] | 4.8060 |