Few-step flow-map generators, including MeanFlow and consistency models, enable efficient sampling through long-range transport, yet their on-policy distillation remains underexplored. We introduce FlowMap-OPD, an on-policy distillation framework that separates student-state acquisition from teacher--student distribution comparison. A formulation based on state marginals establishes this separation, while flow--velocity consistency connects local supervision to the deployed long-range map. Within this framework, we develop flow-map, induced-velocity, and instantaneous-velocity distribution supervision, each paired with a separately specified native flow-map rollout. Cross-capacity ImageNet experiments across three teacher rewards identify instantaneous-velocity distribution supervision with independently tunable student consistency as the most effective choice. In text-to-image experiments, FlowMap-OPD demonstrates strong multi-specialist consolidation capabilities and surpasses multi-reward Flow-Map GRPO in task performance and convergence speed.
Figures & tables
Figure 1: Three specialist teachers, one native student. By separating native rollout from the teacher-supervision kernel, FlowMap-OPD rapidly distills the complementary capabilities of three specialist teachers into a single few-step student: visual preferences, text rendering, and object relations. Large images show outputs from the same student across tasks; the lower-right thumbnails show the corresponding specialist outputs.
Figure 2: From local-transition OPD to few-step flow-map OPD with separate rollout and teacher comparison. (1) Flow/Diffusion OPD [ 4 , 13 ] uses the same local transition distributions from dxt=[vtθ−2gt2∇xlogpt]dt+gtdWt for rollout and teacher–student comparison. (2) FlowMap-OPD collects states through long-range map rollouts, xti+1=ψti→ti+1θ(xti) . (3) Independently of the rollout transition, optimization compares teacher–student distributions through KL(N(μθ,Σ)∥N(μT,Σ)) , with means constructed from direct velocities vθ(x,t) (Section 4.5 ), flow maps ψt→rθ(x) with their stochastic corrections (Section 4.3 ), or induced velocities Vθ,src(x,r,t) and Vθ,dst(x,r,t) (Section 4.4 ).
Figure 3: Specialist capabilities consolidated in the same few-step student. Each row compares the base, all three specialists, and FlowMap-OPD on the same prompt and initial noise with the same few-step sampler. Gold borders mark the specialist for each task; the green column shows the same student across all three tasks: a seed packet labeled “Magic Beans Inside,” a medieval wolf adventurer, and a truck to the left of a refrigerator. Full task scores are reported in Table 2 .
Figure 4: Source and destination consistency. Gray S-curves show the continuous trajectory; dashed arcs show separate long- and short-range flow-map jumps, and black straight arrows show local velocity steps. (a) Source: the direct map ψt→r agrees to first order with a step xt−ε=xt−εvt followed by ψt−ε→r . (b) Destination: ψt→r−ε agrees to first order with ψt→r followed by xr−ε=xr−εvr . Here vt=v(xt,t) and vr=v(xr,r) . These two routes give the identities in Eq. 11 .
Figure 5: Cross-capacity transfer under three teacher rewards. Separate XL/2-to-B/2 runs compare representative supervision recipes. Panels show MMD-setting FID, classifier target log-probability, and DINO cosine similarity under native four-step generation. Every marker is a recorded evaluation; curves stop at the common epoch 300. Dashed lines show the frozen teacher actually used in each sweep. Complete coefficient sweeps and five-step curves appear in Appendix E .
Supervision
MMD
Classifier
DINO
FID ↓
logpy↑
Cosine ↑
B/2 base
22.68
-0.5448
0.7207
XL/2 teacher
14.68
-0.4526
0.8084
Flow-map supervision
Stochastic map, λ=0.1
29.15
-0.9623
0.7004
Stochastic map, λ=0.2
29.15
-0.9621
0.7003
Table 1: Three separate XL/2-to-B/2 reward-transfer settings. All students use epoch 300, with native four-step sampling. Columns report MMD-teacher transfer by FID, classifier transfer by target log-probability, and DINO transfer by cosine similarity. These are different reward settings, not three scores of one student. Best and second-best student results are bold and underlined, respectively.
Task scores
DrawBench scores
Model / supervision
GenEval
OCR
PickScore
PickScore
Aesthetic
DeQA
ImgRwd
UniRwd
Base
0.5041
0.3491
20.9758
21.6298
5.5368
4.1712
0.3918
2.7187
Task-specialized teachers
GenEval teacher
0.8454
—
—
21.8184
5.3905
3.2086
0.6426
2.8719
OCR teacher
—
0.8504
—
21.9869
5.5111
4.1396
0.5386
2.8049
PickScore teacher
—
—
23.0772
23.3711
6.1510
4.0829
1.1402
3.2005
Table 2: Three specialist teachers distilled into one five-step student. The student uses λc=0.001 at update 300. Task scores and DrawBench scores are shown side by side; higher is better.
Figure 6: Three-teacher transfer during training. Full-test-set OCR and PickScore scores during training and at update 300; markers are actual evaluations and lines connect measurements without smoothing. Dashed lines denote the respective specialists. Both panels report the same student with λc=0.001 .
Figure 7: Two specialist teachers versus direct mixed-reward optimization. EMA OCR and PickScore scores for two-teacher OPD and Flow-Map GRPO with equal OCR/PickScore reward weights. Dashed lines denote the task specialists. The shared starting marker is a visual reference.
Figure 8: Consistency (solid, left) versus teacher-map error.(dashed, right) across λc
λc
GenEval ↑
OCR ↑
PickScore ↑
0
0.827
0.830
22.8533
0.001
0.838
0.845
22.9794
0.01
0.725
0.806
22.6529
0.1
0.568
0.646
22.0283
1
0.450
0.340
21.5939
Table 3: MeanFlow consistency sweep at update 200.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Native family
Source interface
Destination interface
Flow-map interface
MeanFlow, two-time map
Self-induced or teacher-directed source predictor; instantaneous-velocity supervision with consistency is also available
Learned destination derivative at the student proposal
Direct flow-map regression or local-anchor Gaussian comparison
CM, endpoint plus re-noising
Endpoint source tangent with declared teacher-velocity closure
Unavailable natively: no learned destination-time derivative
Endpoint-anchor Gaussian comparison or endpoint regression
Appendix
Table 4: Supervision choices supported by each native parameterization.
Figure 9: All twelve supervision configurations for each of the three teacher rewards, under four-step (top) and five-step (bottom) generation. Markers are recorded evaluations, dashed horizontal lines are the frozen teachers, and the dotted vertical line marks the common epoch-300 comparison. MMD FID uses a logarithmic scale; classifier log-probability uses a symmetric logarithmic scale with a linear region between −1 and 1 , retaining unstable configurations. Curves end at the last available evaluation.
MMD: FID ↓
Classifier: logpy↑
DINO: Cosine ↑
Supervision
4 steps
5 steps
4 steps
5 steps
4 steps
5 steps
B/2 base
22.68
22.72
-0.5448
-0.5451
0.7207
0.7215
XL/2 teacher
14.68
14.73
-0.4526
-0.4500
0.8084
0.8096
Stochastic map, λ=0.1
29.15
28.70
-0.9623
-0.9474
0.7004
0.7040
Stochastic map, λ=0.2
29.15
28.70
-0.9621
-0.9472
0.7003
0.7040
Stochastic map, λ=0.4
29.15
28.70
-0.9619
-0.9469
0.7001
0.7037
Appendix
Table 5: Four- and five-step reward-transfer results at epoch 300. Each reward setting reports its corresponding metric. Best student results in each column are bold.
Supervision
Epoch
4 steps
5 steps
8 steps
16 steps
B/2 base
0
22.48
22.72
22.88
22.87
Velocity, λc=0
300
23.58
21.44
19.29
19.17
Velocity, λc=0.01
300
18.14
18.33
18.33
18.56
Velocity, λc=0.1
300
21.59
21.01
20.42
20.41
Velocity, λc=1
300
41.40
40.33
37.19
34.74
Velocity, λc=10
300
395.48
392.19
388.29
386.52
Appendix
Table 6: Fixed epoch-300 checkpoints under the fixed-tail deployment sweep. Lower test FID is better; best student results are bold.
Supervision
Four-step FID ↓
Five-step FID ↓
Stochastic endpoint map, λ=0.1
41.01
39.72
Stochastic endpoint map, λ=0.2
40.91
39.67
Stochastic endpoint map, λ=0.4
40.66
39.28
Stochastic endpoint map, λ=0.8
40.34
38.78
Stochastic endpoint map, λ=1.0
40.12
38.54
Deterministic endpoint map
40.96
39.68
Appendix
Table 7: Supplementary CM distillation at epoch 150. We report FID under four- and five-step generation; lower is better.
λc
Source residual
Velocity mismatch
Average-velocity mismatch
Map composition error
0
0.800
0.017
0.031
0.003391
0.001
0.649
0.020
0.032
0.003148
0.01
0.763
0.053
0.043
0.003504
0.1
0.401
0.088
0.060
0.002558
1
0.179
0.171
0.090
0.001996
Appendix
Table 8: Consistency and teacher-matching errors after 200 training steps, averaged over the preceding 20 steps at the last supervised interval. Velocity errors are mean squared differences from the teacher; map-composition error measures agreement between direct and composed student maps.
Figure 10: Complementary specialist capabilities in one student. Rows compare text rendering, visual preferences and object relations. Columns show the base, OCR teacher, PickScore teacher, GenEval teacher and the same student.
Figure 11: Cross-task specialist–student comparison. Rows cover OCR, visual preferences and compositional generation. Each row compares the base, three specialist teachers and the same student under a shared prompt, initial noise and few-step sampler.
Figure 12: Cross-task specialist–student comparison. Rows cover OCR, visual preferences and compositional generation. Each row compares the base, three specialist teachers and the same student under a shared prompt, initial noise and few-step sampler.
Figure 13: Paired OCR examples. Columns show the base, OCR teacher, PickScore teacher, GenEval teacher, and the same student. Each row shares the prompt, initial noise and few-step sampler.
Figure 14: Paired OCR examples. Columns show the base, OCR teacher, PickScore teacher, GenEval teacher, and the same student. Each row shares the prompt, initial noise and few-step sampler.
Figure 15: Paired PickScore examples. Columns show the base, OCR teacher, PickScore teacher, GenEval teacher, and the same student. Each row shares the prompt, initial noise and few-step sampler.
Figure 16: Paired PickScore examples. Columns show the base, OCR teacher, PickScore teacher, GenEval teacher, and the same student. Each row shares the prompt, initial noise and few-step sampler.
Figure 17: Paired GenEval examples. Columns show the base, OCR teacher, PickScore teacher, GenEval teacher, and the same student. Each row shares the prompt, initial noise and few-step sampler.
Figure 18: Paired GenEval examples. Columns show the base, OCR teacher, PickScore teacher, GenEval teacher, and the same student. Each row shares the prompt, initial noise and few-step sampler.
Department of Computer Science and Engineering, Seoul National University · Department of Computer Science and Engineering, Sogang University Seoul, Republic of Korea