Organizations: College of AI, Tsinghua University · Department of Electrical and Computer Engineering, Princeton University · EPFL · Computer Science Department, Tsinghua University
ELBO-based reinforcement learning offers a sampler-agnostic approach to fine-tuning flow matching models with reward feedback. Timestep weighting in ELBO-based RL has large impact on performance, and it also provides a unified view (as we show in this work) to understand prediction losses heuristically chosen in prior work, yet it remains under-researched and is often chosen to inherit pretrain configs. We investigate impacts and dynamics of timestep weighting in ELBO-based RL. We show that effective weighting depends on both the reward landscape and stage of learning. (1) Through experiments on controlled CIFAR image generation, complemented by robotics, we investigate how weighting impacts reward-driven updates across noise levels. (2) Through gradient analysis, we reveal distinct patterns of cross-noise coordination across tasks and their evolution during training. These findings motivate the hypothesis that useful weighting depends on the gap between the policy's current behavior and the behavior favored by the reward. (3) Guided by this analysis, we study simple static weighting, budgeted profile selection, and dynamic schedules that improve performance beyond conventional target choices. Our results establish timestep weighting as an important design choice for flow-matching RL and motivate further research into methods that choose and adapt it throughout learning.
Figures & tables
Figure 1: CIFAR-10 post-training: mean per-run validation peak under different timestep weighting scheduling (i.e., α : smaller α places more emphasis on larger noise levels) for six image rewards. Intuitively, global rewards prefer timestep weighting that leans toward larger noise levels, but our deeper analysis shows that the preference is also strongly related to the behavioral gap between the model and the reward-favored behavior. More details can be found in Sections 3.2 and 4 , Appendices G.1 and H.1 .
Figure 2: Robot control from scratch: mean per-run peak return under different fixed α for FPO and FPO++, each relative to its native target. Acrobot, Walker and Cartpole prefer more emphasis on larger noise levels than native, while Spot prefers slightly smaller noise. More details can be found in Section 3.2 and Appendices G.2 and H.2 .
Figure 3: Cross-batch gradient cosine between noise bins at an early checkpoint; top/bottom rows are the better/worse α of each task. Post-trained CIFAR rewards show selective agreement among a few bins, while from-scratch Go2 benefits from broad agreement across noise levels. More details can be found in Section 4.1 and Appendix D.2 .
Figure 4: Cross-batch gradient cosine between noise bins over training for fixed- α runs. Post-training (CLIP) gradually shifts learning toward middle and low noise, while from-scratch training (Go2, G1) starts with broad agreement that later weakens. More details can be found in Section 4.2 and Appendix D.2 .
Figure 5: Six CIFAR rewards: mean per-run validation peak of static softmax, compared with target-derived references and fixed- α profiles. Static softmax improves over the best fixed α on CLIP, Mixed and Digit7 and matches it on the others. More details can be found in Appendix E .
Figure 6: Successive halving: peak performance versus training cost for CIFAR (top) and robotics (bottom). At about half the cost of trying all candidates, halving matches or beats the native target in most tasks, but misses Spot’s best profile. More details can be found in Section 5.2 and Appendices E.6 – E.7 .
Figure 7: Dynamic versus fixed timestep weighting on CIFAR (top) and robotics (bottom). CIFAR schedules shift emphasis from larger to smaller noise, and robotics schedules the opposite way; in each task the schedule reaches a higher mean peak than the best fixed α . More details can be found in Section 5.3 and Appendices E.3 and E.4 .
Appendix figures & tables41 assets
Supplementary material from the paper’s appendix.
Appendix
Objective
Profile placement
Residual and reduction
Calibrated CIFAR
Sample t∼qw ; scalar a inside exponent
Velocity error / d ; aggregate, exponentiate, PPO clip
FPO
Explicit wαϵ inside exponent
Noise error; aggregate, exponentiate, PPO clip
FPO++
Explicit wα outside per-draw proxy
Velocity error / d ; process each draw, then average
Appendix
Table 1: Weight placement in the evaluated objectives. Equations ( 6 )–( 7 ), ( 19 ), and ( 20 ) define the operations.
Figure 8: Six CIFAR rewards on a calibrated 7 -exponent by 5 -scale grid ( n=3 per cell): mean per-run validation peak versus α , with one colored curve per whole-loss scale g . Auto-A35, Edge, CLIP, and SimCLR keep their preferred shape across these scales; Mixed’s crossing curves and Digit7’s central switch show where scale can change the winner. Reward axes use native units; the ranks are in Appendix H.1 .
Figure 9: Six CIFAR rewards on the seven-profile, three-seed grid: best mean-peak exponent at each calibrated scale g (left) and mean-peak regret from retaining the g=1 winner (right). Four rewards keep the same winning shape, whereas Mixed and Digit7 incur scale-dependent regret; shape-first tuning may need a recheck for those rewards. The right-panel reward axes are independent; exact selections appear in Table 3 .
Reward
Profile–seed pairs
Peak difference q− explicit
W/T/L
Auto-A35
27
+0.05478±0.15620
13/5/9
Edge
27
+0.00796±0.00538
23/4/0
Appendix
Table 2: Historical matched estimator comparison. Entries are mean paired peak differences ± sample SD across profile–seed pairs; W/T/L uses a .003 tie tolerance. Each reward has nine profiles and three seeds, not 27 independent training seeds.
Figure 10: Five matched CIFAR rewards: initial signed local response and positive opportunity under 4/255 image perturbations (256 pretrained images), compared with the early validation-reward advantage of high-noise α∈{−2,−1} over x0 at updates 20–200 ( n=3 ). Local response varies by reward and does not alone determine the benefit of high-noise emphasis; Mixed’s positive signed response is marked. Native axes are independent; Table 4 defines the response summaries and reports their values.
Figure 11: CIFAR checkpoint diagnostics: within-batch (top) versus independent-cross-batch (bottom) parameter-gradient cosines between clean-to-noisy timestep bins. Apparent within-batch cooperation can weaken across batches, as in successful SimCLR, so shared-sample structure need not be reproducible directional signal; raw-gradient fluctuations exceed stable-mean energy in 75/80 measured states. E/H identify separate training prefixes; color ticks show raw cosines, with sampling and Gram definitions in Appendix D.3 .
Figure 12: Auto-A35 and SimCLR at fixed α=0 : raw training reward trajectories (left) and cross-batch gradient cosine across noise bins at updates 200, 400, and 800 (right). Both deteriorating branches develop stronger cross-bin agreement and larger raw-gradient RMS, so stronger coordination alone does not imply better reward. Colors and axes match Figure 3 ; Appendix D.4 details measurements and the contrasting CLIP branch.
Figure 13: Digit7: cross-batch gradient cosine across noise bins at updates 200/400/800 for fixed α=0,−2,1 ; labels give raw reward and gradient RMS. Reward improves as agreement weakens for 0 , while successful −2 retains common directions: progress has no single cosine signature. Colors show raw-cosine ticks; † flags a ratio departure defined in Appendix D.3 .
Figure 14: CLIP: cross-batch gradient cosine across noise bins at updates 200/400/800 ( α=0 top, −2 bottom). Deteriorating 0 loses agreement, while better −2 retains a middle/low-noise positive block; deterioration has no universal cosine signature. Colors, raw-reward/RMS labels, and the ratio-departure dagger follow Figure 13 .
Figure 15: Six CIFAR rewards: mean per-run fixed-validation peak versus static-softmax concentration ηnorm . Concentration improves some rewards but Digit7 declines above one, so stronger bin selection is not uniformly beneficial. The tested concentration settings are shown on the horizontal axis.
Figure 16: Matched three-seed long-budget FPO cohort: mean of each run’s cumulative evaluation peak through four training-block budgets for fixed αϵ∈{.25,.5,1} . Acrobot’s leading profile changes from .25 at 61 blocks to .5 at 122, back to .25 at 204, and to .5 at 244; Walker retains 1>.5>.25 throughout. Lines connect budgets, not changing-weight schedules. Each seed contributes a single run, so values can differ from Figure 2 , which averages repeated runs per seed; the αϵ=0 arm is reported in the text.
Figure 17: V-GRPO image rewards versus timestep exponent α : held-out observed reward from one training seed per arm. Preferred exponents differ by reward, extending the task-dependent weighting pattern beyond CIFAR; the displayed values are descriptive rather than replicated comparisons. OCR/TextPecker use peaks over updates 5–25, GenEval the update-50 value, and PickScore peaks over updates 100–300; trajectories are in Figure 18 .
Figure 18: One-seed V-GRPO held-out evaluations: TextPecker across tested exponents (left), and PickScore trajectories at updates 100, 200, and 300 (right). TextPecker’s α=2 arm collapses; PickScore’s −.5 and velocity arms nearly tie in observed peak while negative exponents retain stronger update-300 scores. These checkpoints explain the peak summaries in Figure 17 without using training maxima.
Figure 19: Six CIFAR rewards: within-scale ranks of seven α profiles across five calibrated whole-loss scales, computed from mean per-run validation peaks ( n=3 ). Most rewards retain similar shape rankings across scales, whereas Mixed and Digit7 change leading profiles, qualifying shape-first tuning. Colors share a rank scale rather than reward units; native reward surfaces are shown in Figure 8 .
Reward
g=.5
.75
1
1.5
2
Max regret
Auto-A35
−1
−1
−1
−1
−1
0.0000
Mixed
−1
−1
−1
−0.5
−0.5
0.1268
Edge
0
0
0
0
0
0.0000
CLIP
−1
−1
−1
−1
−1
0.0000
SimCLR
−1
−1
−1
−1
−1
0.0000
Digit7
0
−0.5
−0.5
0
0
0.0248
Appendix
Table 3: Selected shape at each scale and maximum regret of retaining the g=1 winner. Winners maximize mean per-seed validation peak on the seven-profile grid; regret is in each reward’s native units. These are descriptive grid selections, not an independent transfer experiment.
Reward
DD (%)
M
C
O
Explicit Δ
q cohort Δ
Automobile
49.87
.005621
−.005213
.000733
−.011499
.108042
Auto-A35
48.70
.004578
−.004230
.000638
.018274
.128433
CLIP
36.85
.010949
−.003178
.009414
.035520
.091558
SimCLR
40.76
.016113
−.005121
.015405
.027671
.052726
Digit7
46.35
.008746
−.001750
.008322
.002082
—
Edge
6.25
.000209
.010453
.011644
−.002609
−.003045
Appendix
Table 4: Initial fine response at amplitude 4/255 and mean reward advantage over updates 20–200. M,C,O follow Appendix D.1 ; DD is double-decrease percentage. The last two columns compare the fixed {−2,−1} average with the same-cohort x0 control. Explicit and q -sampling cohorts remain separate.
Statistic
Value
Digit7 q , updates 20–40
−.004286
Digit7 q , updates 20–80
−.002453
Digit7 q , updates 20–200
.014318
Explicit radius 4, H=200 : (n,ρ,p)
(6,.60,.24)
Leave-one-reward ρ , explicit
[.40,.90]
Leave-one-reward ρ , q cohort
[−.30,.70]
Appendix
Table 5: Multiscale response summaries. The Digit7 rows report mean advantage of the negative-profile pair over x0 ; the remaining rows are cross-reward descriptive associations.
Cohort
Radius
H=40
80
160
200
Explicit
2
−.20
−.20
−.03
.60
Explicit
4
−.20
−.20
−.03
.60
Explicit
8
.03
.03
−.37
.14
q cohort
2
.83
.14
.14
.26
q cohort
4
.83
.14
.14
.26
q cohort
8
.89
−.03
−.03
.09
Appendix
Table 6: Descriptive cross-reward Spearman associations of standardized fine-damage magnitude with standardized early mean advantage, across six rewards within each cohort. Both quantities are divided by the initial unperturbed reward SD. Radius is in units of 1/255 .
Task
α
Checkpoint
Training seeds
Peak reward/return
SimCLR
−2/1
200
20260724 / 20260724
.999091/.667570
CLIP
−2/1
200
20260724 / 20260724
.895339/.250209
Digit7
−1/1
200
20260724 / 20260726
.621488/.285237
Go2
1/0
300
9200, 9201 (both rows)
41.590/39.385
Appendix
Table 7: Profiles and full-trajectory peaks underlying Figure 3 . Paired entries are higher/lower rows. CIFAR checkpoints are optimizer updates and Go2 checkpoints are training iterations.
α
Training seed
Update
Training reward
Cross-bin mean
-2.0000
20260725
100
0.3472
0.1448
-1.0000
20260724
100
0.3368
0.1520
0.0000
20260724
100
0.3353
0.1787
1.0000
20260726
100
0.2886
0.0778
Appendix
Table 8: Digit7 at the independent early-prefix checkpoint. Each row is one selected training seed; entries report recorded raw training reward and mean off-diagonal cross-batch cosine. The better-reward branches have stronger agreement than the velocity branch at this checkpoint. These E-prefix states are not joined to the H-prefix trajectories in Figures 3 and 4 .
Task
High–high
Low–low
High–low
Walker
.537
.375
−.311
Swimmer
.400
.723
−.380
Go2
.395
.275
−.182
Appendix
Table 9: Native-exponent late-phase cross-batch cosine by noise block. High noise is σ≥.6 and low noise is σ≤.3 ; within-block means omit diagonals. Cross-block means include both symmetric rectangles.
Method
Task
Checkpoints
Median within–cross gap
FPO
Acrobot
24
0.0724
FPO
Ball in Cup
23
0.1050
FPO
Cheetah
22
0.0668
FPO
Fish
24
0.1055
FPO
Swimmer
24
0.0461
FPO
Walker
24
0.0453
Appendix
Table 10: Checkpoint inventory and median off-diagonal within–cross cosine gap, by task. Every checkpoint gap is positive.
Quantity
All 80
Omit two †
tr(N)>tr(K)
75/80
73/78
Mean off-diagonal Cfluct>0
80/80
78/78
Median S=tr(K)/tr(N)
0.17480
0.15766
Median mean off-diagonal Cfluct
0.09156
0.09335
Median fraction of positive off-diagonal entries
80.30%
80.30%
States with negative off-diagonal tenth percentile
68/80
66/78
Appendix
Table 11: Shared-fluctuation summaries over the 80 measured states and after omitting the two ratio-marked states. Correlation summaries first average entries within a state, then take the median across states. State counts are descriptive, not independent training replications.
Reward
States
Median S
Median mean Cfluct
Auto-A35
20
0.09546
0.09064
CLIP
20
0.19128
0.09596
SimCLR
20
0.08243
0.08937
Digit7
20
0.40711
0.09098
Appendix
Table 12: Per-reward medians of the energy ratio S and mean off-diagonal fluctuation correlation. These summarize checkpoint geometry, not a reward-quality ordering.
Reward
Rule
Peak validation reward
Auto-A35
Static direct
0.968577±0.039152
Auto-A35
Online coherence
0.844915±0.259461
Mixed
Static direct
0.672966±0.279437
Mixed
Online coherence
0.542469±0.175938
Edge
Static direct
0.967028±0.002040
Edge
Online coherence
0.960761±0.008673
Appendix
Table 13: Additional fixed-scale methods: mean per-seed validation peak ± sample SD ( n=3 ). The shared fixed controls and selected static-softmax results are in Figure 5 ; Digit7’s distinct matched method-cohort x0 control is retained here. Online starts uniform except in the explicitly labeled Digit7 row.
Reward
ηraw
Peak validation reward
Updates
Auto-A35
.5
.3024±.0518
200–280
CLIP
.25
.2316±.0168
200–440
Digit7
1
.3161±.0428
200–300
Appendix
Table 14: Best concentration per reward for full-refit online weighting, selected by mean per-seed validation peak. Uncertainty is sample SD, n=3 . Every run at these selected settings stopped by degradation; update ranges describe these selected settings only.
Reward
Allocation
Observed peak
Step 800
CLIP
Fixed −2
0.852268±0.027500
—
CLIP
Fixed −3
0.864639±0.012958
—
CLIP
Fixed 0
0.387094±0.023369
—
CLIP
−3→0
0.910677±0.011321
0.907311±0.010997
CLIP
Static softmax η=1.25
0.899151±0.006857
—
Digit7
Fixed +1
0.311454±0.042482
—
Appendix
Table 15: Three-reward path interventions and controls: mean per-seed validation peak and step-800 reward ± sample SD, three seeds. A dash denotes the absence of a complete three-seed step-800 control, not a zero reward. Peaks use every trajectory’s observed evaluations. The Digit7 static-method cohort in Figure 5 is separate.
Task
Native
Selected
Oracle
Δ
Gap
Acrobot
163.631157
185.187285
191.103217
+21.556128
5.915933
Spot
313.813333
233.430000
313.813333
-80.383333
80.383333
Walker
724.802788
742.842999
754.124350
+18.040210
11.281352
Appendix
Table 16: Robotics replay in Figure 6 , bottom: n=3 , first cut 20% . FPO uses the 244-block seed-901–903 cohort with repeated runs averaged pointwise within seed; Spot retains its 1,500-update cohort. Selected is the expected empirical mean-peak score, oracle is the best fixed-reference score, Δ is selected minus native, and gap is oracle minus selected. The lower block gives selection probabilities, the oracle exponent, and search/exhaustive fractions in native training budgets. Exponents are αϵ for FPO and α for Spot.
Task
First cut
Selected score
Exact draw range
C
Acrobot
10%
182.14
[176.86, 191.10]
4.1803
Acrobot
20%
185.19
[176.86, 191.10]
5.3607
Acrobot
40%
188.88
[176.86, 191.10]
7.7582
Acrobot
Full search
187.30
[176.86, 191.10]
12.0000
Spot
10%
197.73
[197.73, 197.73]
3.8060
Spot
20%
233.43
[233.43, 233.43]
4.8060
Appendix
Table 17: Sensitivity to first-cut budget in the recorded replay, n=3 (244-block FPO cohort; unchanged 1,500-update Spot cohort). The range is the exact minimum/maximum across bootstrap selection draws, not a confidence interval over independently repeated training. Full search still selects from resampled seeds; its expected selected score therefore need not equal the fixed-reference oracle.
Figure 20: Acrobot fixed-profile robot policies: phase-averaged cross-batch gradient cosines across clean-to-noisy bins (columns: early, middle, late; rows: tested exponents). Early cross-noise agreement is stronger at αϵ=0 than at 1, but the former weakens by the late phase. Colors show the common raw-cosine scale, and n counts contributing seeds; Appendix D.5 defines phase aggregation.
Figure 21: Ball in Cup fixed-profile robot policies: phase-averaged cross-batch gradient cosines across clean-to-noisy bins (columns: early, middle, late; rows: tested exponents). The early αϵ=2 panels show more agreement than its middle phase, illustrating stage-dependent rather than fixed coordination. Colors show the common raw-cosine scale, and n counts contributing seeds; Appendix D.5 defines phase aggregation. At exponent 2, seed 901 failed at block 114; the late panel contains seed 902 only.
Figure 22: Cheetah fixed-profile robot policies: phase-averaged cross-batch gradient cosines across clean-to-noisy bins (columns: early, middle, late; rows: tested exponents). Early αϵ=0 agreement weakens late, while the 2 panels remain near zero on average. Colors show the common raw-cosine scale, and n counts contributing seeds; Appendix D.5 defines phase aggregation. At exponent 2, seed 902 failed at block 58; middle and late panels contain seed 901 only.
Figure 23: Fish fixed-profile robot policies: phase-averaged cross-batch gradient cosines across clean-to-noisy bins (columns: early, middle, late; rows: tested exponents). Cross-noise agreement declines from early to late at every displayed exponent. Colors show the common raw-cosine scale, and n counts contributing seeds; Appendix D.5 defines phase aggregation.
Figure 24: Swimmer fixed-profile robot policies: phase-averaged cross-batch gradient cosines across clean-to-noisy bins (columns: early, middle, late; rows: tested exponents). The αϵ=.5 late panel retains more agreement than its early panel, unlike several other profiles. Colors show the common raw-cosine scale, and n counts contributing seeds; Appendix D.5 defines phase aggregation.
Figure 25: Walker fixed-profile robot policies: phase-averaged cross-batch gradient cosines across clean-to-noisy bins (columns: early, middle, late; rows: tested exponents). The αϵ=0 late panel gains agreement, whereas the native 1 panels also contain opposing noise blocks. Colors show the common raw-cosine scale, and n counts contributing seeds; Appendix D.5 defines phase aggregation.
Figure 26: Cartpole fixed-profile robot policies: phase-averaged cross-batch gradient cosines across clean-to-noisy bins (columns: early, middle, late; rows: tested exponents). Cross-noise agreement remains broadly high for every recorded profile, even as its magnitude changes by phase. Colors show the common raw-cosine scale, and n counts contributing seeds; Appendix D.5 defines phase aggregation.
Figure 27: Cartpole fixed-profile robot policies: phase-averaged cross-batch gradient cosines across clean-to-noisy bins (columns: early, middle, late; rows: tested exponents). Cross-noise agreement remains broadly high for every recorded profile, even as its magnitude changes by phase. Colors show the common raw-cosine scale, and n counts contributing seeds; Appendix D.5 defines phase aggregation.
Figure 28: G1 fixed-profile robot policies: phase-averaged cross-batch gradient cosines across clean-to-noisy bins (columns: early, middle, late; rows: tested exponents). Off-diagonal agreement stays weak across all recorded profiles and phases, contrasting with Cartpole. Colors show the common raw-cosine scale, and n counts contributing seeds; Appendix D.5 defines phase aggregation.
Figure 29: Go2 fixed-profile robot policies: phase-averaged cross-batch gradient cosines across clean-to-noisy bins (columns: early, middle, late; rows: tested exponents). The α=0 late phase retains more overall agreement than the negative-exponent profiles despite their stronger early agreement. Colors show the common raw-cosine scale, and n counts contributing seeds; Appendix D.5 defines phase aggregation.
Figure 30: Go2 fixed-profile robot policies: phase-averaged cross-batch gradient cosines across clean-to-noisy bins (columns: early, middle, late; rows: tested exponents). The native α=1 profile loses much of its early agreement by the late phase; its higher peak return than α=0 does not imply stronger late coordination. Colors show the common raw-cosine scale, and n counts contributing seeds; Appendix D.5 defines phase aggregation.
Figure 31: Spot fixed-profile robot policies: phase-averaged cross-batch gradient cosines across clean-to-noisy bins (columns: early, middle, late; rows: tested exponents). The native α=1 late phase retains more agreement than α=0 , illustrating a different within-task trajectory from Go2. Colors show the common raw-cosine scale, and n counts contributing seeds; Appendix D.5 defines phase aggregation.