CapField-OPD: Learning Continuous Capability Fields via Joint-Anchored Multi-Teacher On-Policy Distillation for Flow Models
Authors: Pengyang Ling, Jiazi Bu, Yujie Zhou, Yibin Wang, Zeqiang Lai, Xiaoxiao Ma, Yi Jin, Huaian Chen, +1 more
Organizations: University of Science and Technology of China · Shanghai Jiao Tong University · Fudan University · The Chinese University of Hong Kong · Shanghai Artificial Intelligence Laboratory
Reward-specialized post-training produces strong experts for flow-based generative models, while multi-teacher on-policy distillation (OPD) consolidates their capabilities into a single student. Existing methods, however, route each prompt to a single teacher according to its semantic category, implicitly binding the desired capability to prompt content. This coupling makes capability invocation vulnerable to prompt perturbations and prevents users from explicitly adjusting the strength of the desired capability at inference time. In this work, we introduce CapField-OPD, an OPD framework that integrates multiple teachers into a continuous capability field through explicit capability coordinates. We use teacher models as anchors to construct this field, with the coordinates determining how their outputs are combined. Each capability configuration thus receives a unique supervision target, and capability control no longer depends on prompt semantics. Since the training anchors may not be optimal at inference time, we further profile the learned field on a small calibration set. The coordinate with the highest mean reward serves as the recommended default, while coordinates that are frequently optimal offer a promising candidate set for test-time scaling. Extensive experiments on compositional generation, text rendering, and visual aesthetics demonstrate that CapField-OPD consolidates multiple specialized teachers into a single student while preserving or surpassing their performance, reliably invokes the desired capabilities under semantics-preserving prompt variations, and supports continuous capability control and coordinate-based test-time scaling.
Figures & tables
Figure 1: Overview of CapField-OPD. (a) Teacher anchors define a continuous capability field controlled by explicit, potentially extrapolative coordinates. (b-d) The learned field supports continuous control, robust capability invocation under prompt rewriting, and coordinate-based test-time scaling.
Figure 2: Visualization under prompt rewriting.
Method
Composition-task prompts
Text-rendering prompts
Aesthetic prompts
GenEval
HPSv3
CLIP
PickScore
OCR
HPSv3
CLIP
PickScore
HPSv3
CLIP
PickScore
FLUX.1-dev
0.6616
8.65
0.3966
23.40
0.5735
13.00
0.4466
22.91
13.28
0.3868
22.58
GenEval teacher
0.9417
9.16
0.4092
23.27
0.6170
13.43
0.4585
22.97
13.49
0.3946
22.47
OCR teacher
0.7166
8.90
0.4030
23.59
0.9417
12.74
0.4567
22.88
13.57
0.3834
22.64
Aesthetic teacher
0.3228
10.62
0.4117
24.04
0.4824
15.14
0.4688
23.95
15.23
0.4181
23.68
GenEval+Aesthetic teacher
0.9253
11.75
0.4210
24.11
0.6235
14.74
0.4525
23.32
14.65
0.3968
22.71
Table 1: Quantitative comparison. Single-only and Single+Joint distill from the three single-capability teachers without and with the two joint-capability teachers, respectively. The two CapField-OPD modes share one student model and differ only in their inference coordinates. Bold and underlined values denote the best and second-best results among unified models. Since aesthetic prompts contain no composition target or text-rendering target, the joint mode does not apply, and the corresponding entries are marked “–”.
Figure 3: Visual demonstration of different methods, from left to right: FLUX.1-dev, Multi-task GRPO, DiffusionOPD(Single-only), DiffusionOPD(Single+Joint), DanceOPD(Single-only), DanceOPD(Single+Joint), CapField-OPD(Single mode), and CapField-OPD(Joint mode).
Metric
FLUX.1-dev
Multi-task GRPO
DiffusionOPD
DanceOPD
CapField-OPD Single mode
Single
Single+Joint
Single
Single+Joint
GenEval ↑
0.6550
0.8461
0.6241
0.6132
0.6838
0.5560
0.9267
OCR ↑
0.6115
0.9242
0.8662
0.8718
0.8879
0.8097
0.9677
HPSv3 ↑
13.46
14.45
15.13
15.15
15.08
15.17
15.28
Table 2: Robustness evaluation under semantics-preserving prompt rewriting.
Figure 4: Illustration of continuous capability control at inference time.
Figure 5: Demonstration of coordinate-based and seed-based test-time scaling.
Training anchors
Composition + Aesthetics
Text Rendering + Aesthetics
GenEval ↑
HPSv3 ↑
CLIP ↑
PickScore ↑
OCR ↑
HPSv3 ↑
CLIP ↑
PickScore ↑
Base + single
0.7723
9.38
0.4034
22.23
0.8716
13.50
0.4481
22.84
Base + single + joint
0.9322
11.76
0.4201
23.94
0.9256
14.08
0.4538
23.32
Table 3: Effect of joint-teacher anchors at jointly activated tasks.
Coordinate selection
Composition
Text Rendering
Aesthetics
λ^peakS
GenEval ↑
λ^peakS
OCR ↑
λ^peakS
HPSv3 ↑
CLIP ↑
PickScore ↑
Teacher model
/
0.9417
/
0.9417
/
15.23
0.4181
23.68
In-range profiling
(1.0, 0.0, 0.0)
0.9487
(0.0, 1.0, 0.0)
0.9390
(0.0, 0.0, 1.0)
15.25
0.4171
23.70
Expanded profiling
(1.05, 0.0, 0.0)
0.9531
(0.0, 1.5, 0.0)
0.9592
(0.0, 0.0, 1.3)
15.32
0.4182
23.71
Table 4: Effect of coordinate extrapolation. Each λ^peakS (Eq. 11 ) reports the full selected coordinate over (GenEval, OCR, aesthetics), with S being the single capability of the corresponding metric.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Value
Parameter
Value
Random seed
42
Learning rate
2×10−4
Optimizer
AdamW
Weight decay
1×10−4
AdamW β1,β2
0.9,0.999
AdamW ϵ
1×10−8
LR scheduler
Constant
Text length (CLIP / T5)
77 / 256
LoRA rank / α
64 / 128
Cond. dimensions
3→256→3072
Batch size
8
Number of GPUs
8
Appendix
Table 5: Hyperparameter settings used for CapField-OPD.
Figure 6: Capability curves during CapField-OPD training.
Figure 7: Additional coordinate sweeps at inference time from GenEval to aesthetics.
Figure 8: Additional coordinate sweeps at inference time from OCR to aesthetics.