Task vectors provide a simple mechanism for composing learned capabilities through model merging. However, the composability of task vectors produced by on-policy distillation (OPD) remains largely unexplored. OPD trains a student using teacher feedback on student-generated trajectories, yielding parameter updates that differ from those produced by the teacher model, usually by reinforcement learning (RL). We therefore ask whether OPD task vectors can complement their RL teacher updates and compose effectively across tasks. Across five domains and two model architectures, we find evidence for both forms of composability. Within a task, merging OPD and RL task vectors can outperform both constituent models, even when the OPD student is weaker than its RL teacher. Across tasks, OPD task-vector compositions achieve higher average scores than corresponding RL compositions in seven of eight backbone-merging-rule comparisons. Parameter-space analyses reveal substantial non-collinearity between OPD and RL updates. Experiment in CODE domain on SMOLLM3-3B shows that the combined direction outperforms either constituent direction at the tested global update norm, supporting directional complementarity in this configuration. Across tasks, OPD updates also show lower overlap among the top-10% feed-forward channels ranked by update energy. Together, these results show that weaker standalone performance does not imply weaker task-vector composability. OPD task vectors can complement stronger RL teacher updates and combine effectively across tasks, highlighting composability as a distinct property for understanding and evaluating post-training updates.
Figures & tables
Figure 1: OPD task-vector composition on Qwen3-4B . (a) Best within-task OPD–RL merges; labels show gains over RL (pp). (b) Specialists (dashed) and TIES merges (solid); radial scores are min–max normalized separately for each task.
Method
Math
Code
IF
RL
60.14
37.51
69.53
OPD
59.58
37.18
68.60
TA
60.51 (+0.37)
37.48 (-0.03)
68.92 (-0.61)
TIES
60.35 (+0.21)
41.24 (+3.73)
69.66 (+0.13)
TSV-M
56.51 (-3.63)
38.52 (+1.01)
69.29 (-0.24)
Raw Sum
59.78 (-0.36)
41.01 (+3.50)
70.83 (+1.30)
Table 1: Within-task composition on Qwen3-4B (%).
SmolLM3-3B
Qwen3-4B
Method
Src.
Math
Code
IF
Avg.
Math
Code
IF
Sci.
Agent
Avg.
Single
RL
24.38
23.08
48.23
31.90
60.14
37.51
69.53
44.95
74.00
57.23
OPD
20.42
22.69
47.92
30.34
59.58
37.18
68.60
42.42
73.00
56.16
Raw Sum
RL
25.21
23.39
48.75
32.45
58.19
40.97
69.51
35.86
69.50
54.81
OPD
24.79
24.24
47.14
32.06
60.56
40.04
68.71
38.89
70.00
55.64
TA
RL
18.31
18.23
44.14
26.89
53.61
34.87
58.18
42.93
64.50
50.82
Table 2: Cross-task composition of OPD and RL task vectors (%). Avg.: equal-weight domain mean. Bold marks the higher displayed score in each merging pair (both for ties), not statistical significance. Single denotes separate task-specific experts.
SmolLM3-3B
Qwen3-4B
Method
Math
Code
IF
Avg.
Math
Code
IF
Sci.
Agent
Avg.
RL
25.21
23.39
48.75
32.45
58.19
40.97
69.51
35.86
69.50
54.81
OPD
24.79
24.24
47.14
32.06
60.56
40.04
68.71
38.89
70.00
55.64
OPD + RL
26.04
24.64
49.16
33.28
60.42
40.39
68.64
40.40
73.50
56.67
Table 3: Joint composition of OPD and RL task vectors across backbones. Scores are percentages; Avg. equally weights the domains available for each backbone. Bold marks column-wise maxima within each backbone.
Method
Math
Code
IF
Avg.
RL
25.21
23.39
48.75
32.45
OPD
24.79
24.24
47.14
32.06
OPD + RL
26.04
24.64
49.16
33.28
Naive M-OPD
21.26
19.26
43.64
28.05
Open-MOPD
22.42
21.73
49.58
31.24
Table 4: Comparison with multi-model OPD baselines on SmolLM3-3B . Scores are percentages; Avg. equally weights Math, Code, and IF. Bold marks column-wise maxima.
Figure 2: SmolLM3-3B task-vector structure. (a) Code layer-wise energy share; shading marks the final 12 layers and Δ denotes OPD minus RL (pp). (b) Update L2 norms and (c) BF16-visible nonzero fractions for Math and Code. (d) Code rank- k energy retention F(k) . Full profiles: Fig. 5 .
Figure 3: Layer-wise RL–OPD alignment on SmolLM3-3B. Task-vector matrix cosine similarities for (a) Math and (b) Code, with a shared color scale. Columns index Transformer layers; rows index attention (Q, K, V, O) and MLP (Gate, Up, Down) projections.
Figure 4: Cross-task overlap and direction retention on Qwen3-4B . (a,b) Jaccard overlap of RL/OPD top- 10% MLP-channel sets (shared scale). (c) OPD–RL difference in ct(λ)=cos(ut,ut+λvt) ; positive values favor OPD, λ=1 denotes Raw Sum.
Table 6: Training-data provenance and released split sizes. Each row supplies the prompts for one domain-specific RL expert and its corresponding OPD student.
Domain
Benchmark
Instances
Metric
Backbones
Math
AIME 2024
30
Answer accuracy, avg@8
Both
AIME 2025
30
Answer accuracy, avg@8
Both
Code
LiveCodeBench v5
167
Execution correctness, avg@10
Both
LiveCodeBench v6
175
Execution correctness, avg@10
Both
IF
IFEval
541
Strict prompt-level accuracy
Both
IFBench (test)
300
Strict prompt-level accuracy
Both
Appendix
Table 7: Evaluation datasets and per-benchmark metrics. Counts are benchmark instances before response sampling and averaging across experimental repetitions.
SmolLM3-3B
Qwen3-4B
Hyperparameter
Math
Code
IF
All five domains
Optimizer
AdamW
AdamW
AdamW
AdamW
Learning rate
10−6
10−6
10−6
10−6
Training batch size
128
128
128
128
Mini-batch size
32
32
32
128
Rollout n
4
4
4
4
Appendix
Table 8: OPD training configurations. Rollout n is the number of responses generated per training prompt.
Figure 5: Complete parameter-structure profiles for Math and Code. (a) Layer-wise squared L2 norms normalized by full-vector energy, with the final 12 Transformer layers shaded; lower strips show OPD minus RL in percentage points. (b) Task-vector L2 norms (top) and nonzero-coordinate fractions in task vectors computed from saved BF16 weights (bottom). (c) Spectral concentration F(k) , estimated as the fraction of total matrix-update energy captured by rank- k approximations across 252 attention and MLP projection matrices. The six measured ranks are 8, 16, 32, 64, 128, and 256, shown on a logarithmic axis. These are descriptive statistics of the evaluated checkpoints.
On-policy distillation (OPD) trains a student on its own responses with token-level feedback from a stronger teacher, yet the teacher can fail on questions the student already answers correctly, and how often each model succeeds varies across tasks. OPD ignores these outcomes and, on average, pushes down even the student's correct responses; gating feedback by student correctness fixes the direction but uses the teacher in the same way whether or not it succeeded. We introduce DuoOPD, in which the student's outcome sets the direction of feedback and the joint teacher-student outcome decides how the teacher supports it: when only the teacher succeeds, its verified answer becomes context for scoring the student's failed response, and when only the student succeeds, a weight shared within the task reinforces the whole response. A single rule covers all four outcome combinations without task-specific settings. Across Qwen3 and Llama, DuoOPD outperforms all five baselines in mean macro accuracy, improving over OPD by 2.58 and 5.98 percentage points, and it also leads on two further task mixtures spanning scientific calculation, instruction following, and code generation. Ablations show that outcome-based direction alone stays near the gated baseline, while the joint-outcome designs supply most of the gain.
Ao Yu, Weibo Gao, Heng Zhou +6
University of Science and Technology of China · The Hong Kong Polytechnic University · The University of Hong Kong +1
Multi-teacher on-policy distillation (MOPD) has emerged as a popular post-training paradigm for integrating specialized capabilities in frontier language models. Existing OPD research has primarily focused on optimizing single-task distillation through objective design, distillation scope, and teacher signal construction, whereas MOPD must aggregate multiple capabilities in shared parameters and address the resulting capability seesaw, in which improving one domain suppresses capabilities acquired from another. Inspired by the distinctive update geometry of OPD, we find that parameter updates from different tasks rapidly concentrate in their respective low-dimensional subspaces during MOPD, providing a direct geometric basis for identifying and controlling cross-task interference. We therefore propose PMOPD (Projection-based Multi-Teacher On-Policy Distillation), which constructs subspace memories from the cumulative parameter displacements of different tasks and projects both gradients and optimizer updates to remove components that interfere with protected task directions. We further develop a lightweight conflict probe to characterize task interactions and guide task ordering, together with a cycling strategy that balances subspace estimation and timely task revisitation. Experiments on representative Code, Reason, and Math tasks show that PMOPD improves every evaluated capability over MOPD, raising the average score across the three tasks by 2.54 points on Qwen2.5-7B and 2.09 points on Llama-3.1-8B. These consistent gains establish geometry-aware optimization as an effective and transferable approach to balanced multi-teacher distillation.
On-policy distillation (OPD) provides dense token-level supervision on student-generated trajectories. Although OPD performs strongly when teacher and student belong to the same model family, we find that its effectiveness degrades in cross-family settings even after tokenizer alignment, with substantially stronger external teachers offering little additional improvement. To understand this disconnect, we decompose the cross-family OPD signal into two components: an offset between a low-capability teacher-family reference and the student, and the within-family log-likelihood shift from that reference to the strong teacher. Standard OPD transfers both components together, allowing the offset to dominate the update direction and obscure the changes associated with teacher capability improvements. We propose CompassOPD, which removes this offset and transfers the within-family shift, while a frozen student reference anchors updates to the student's initial policy. Thus, both teacher-side and student-side changes are measured within their respective model families. Experiments across three student families and multiple teacher families show that CompassOPD consistently outperforms standard cross-family OPD, improving average reasoning accuracy by up to 5.50 points. For an MoE teacher, we further construct the reference directly from the teacher checkpoint by reducing expert activation, eliminating the need for a separate reference checkpoint while retaining a 3.43-point gain over OPD.
Naibin Gu, Qingyi Si, Chenxu Yang +5
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China · School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China