Task vectors provide a simple mechanism for composing learned capabilities through model merging. However, the composability of task vectors produced by on-policy distillation (OPD) remains largely unexplored. OPD trains a student using teacher feedback on student-generated trajectories, yielding parameter updates that differ from those produced by the teacher model, usually by reinforcement learning (RL). We therefore ask whether OPD task vectors can complement their RL teacher updates and compose effectively across tasks. Across five domains and two model architectures, we find evidence for both forms of composability. Within a task, merging OPD and RL task vectors can outperform both constituent models, even when the OPD student is weaker than its RL teacher. Across tasks, OPD task-vector compositions achieve higher average scores than corresponding RL compositions in seven of eight backbone-merging-rule comparisons. Parameter-space analyses reveal substantial non-collinearity between OPD and RL updates. Experiment in CODE domain on SMOLLM3-3B shows that the combined direction outperforms either constituent direction at the tested global update norm, supporting directional complementarity in this configuration. Across tasks, OPD updates also show lower overlap among the top-10% feed-forward channels ranked by update energy. Together, these results show that weaker standalone performance does not imply weaker task-vector composability. OPD task vectors can complement stronger RL teacher updates and combine effectively across tasks, highlighting composability as a distinct property for understanding and evaluating post-training updates.
Figures & tables
Figure 1: OPD task-vector composition on Qwen3-4B . (a) Best within-task OPD–RL merges; labels show gains over RL (pp). (b) Specialists (dashed) and TIES merges (solid); radial scores are min–max normalized separately for each task.
Method
Math
Code
IF
RL
60.14
37.51
69.53
OPD
59.58
37.18
68.60
TA
60.51 (+0.37)
37.48 (-0.03)
68.92 (-0.61)
TIES
60.35 (+0.21)
41.24 (+3.73)
69.66 (+0.13)
TSV-M
56.51 (-3.63)
38.52 (+1.01)
69.29 (-0.24)
Raw Sum
59.78 (-0.36)
41.01 (+3.50)
70.83 (+1.30)
Table 1: Within-task composition on Qwen3-4B (%).
SmolLM3-3B
Qwen3-4B
Method
Src.
Math
Code
IF
Avg.
Math
Code
IF
Sci.
Agent
Avg.
Single
RL
24.38
23.08
48.23
31.90
60.14
37.51
69.53
44.95
74.00
57.23
OPD
20.42
22.69
47.92
30.34
59.58
37.18
68.60
42.42
73.00
56.16
Raw Sum
RL
25.21
23.39
48.75
32.45
58.19
40.97
69.51
35.86
69.50
54.81
OPD
24.79
24.24
47.14
32.06
60.56
40.04
68.71
38.89
70.00
55.64
TA
RL
18.31
18.23
44.14
26.89
53.61
34.87
58.18
42.93
64.50
50.82
Table 2: Cross-task composition of OPD and RL task vectors (%). Avg.: equal-weight domain mean. Bold marks the higher displayed score in each merging pair (both for ties), not statistical significance. Single denotes separate task-specific experts.
SmolLM3-3B
Qwen3-4B
Method
Math
Code
IF
Avg.
Math
Code
IF
Sci.
Agent
Avg.
RL
25.21
23.39
48.75
32.45
58.19
40.97
69.51
35.86
69.50
54.81
OPD
24.79
24.24
47.14
32.06
60.56
40.04
68.71
38.89
70.00
55.64
OPD + RL
26.04
24.64
49.16
33.28
60.42
40.39
68.64
40.40
73.50
56.67
Table 3: Joint composition of OPD and RL task vectors across backbones. Scores are percentages; Avg. equally weights the domains available for each backbone. Bold marks column-wise maxima within each backbone.
Method
Math
Code
IF
Avg.
RL
25.21
23.39
48.75
32.45
OPD
24.79
24.24
47.14
32.06
OPD + RL
26.04
24.64
49.16
33.28
Naive M-OPD
21.26
19.26
43.64
28.05
Open-MOPD
22.42
21.73
49.58
31.24
Table 4: Comparison with multi-model OPD baselines on SmolLM3-3B . Scores are percentages; Avg. equally weights Math, Code, and IF. Bold marks column-wise maxima.
Figure 2: SmolLM3-3B task-vector structure. (a) Code layer-wise energy share; shading marks the final 12 layers and Δ denotes OPD minus RL (pp). (b) Update L2 norms and (c) BF16-visible nonzero fractions for Math and Code. (d) Code rank- k energy retention F(k) . Full profiles: Fig. 5 .
Figure 3: Layer-wise RL–OPD alignment on SmolLM3-3B. Task-vector matrix cosine similarities for (a) Math and (b) Code, with a shared color scale. Columns index Transformer layers; rows index attention (Q, K, V, O) and MLP (Gate, Up, Down) projections.
Figure 4: Cross-task overlap and direction retention on Qwen3-4B . (a,b) Jaccard overlap of RL/OPD top- 10% MLP-channel sets (shared scale). (c) OPD–RL difference in ct(λ)=cos(ut,ut+λvt) ; positive values favor OPD, λ=1 denotes Raw Sum.
Table 6: Training-data provenance and released split sizes. Each row supplies the prompts for one domain-specific RL expert and its corresponding OPD student.
Domain
Benchmark
Instances
Metric
Backbones
Math
AIME 2024
30
Answer accuracy, avg@8
Both
AIME 2025
30
Answer accuracy, avg@8
Both
Code
LiveCodeBench v5
167
Execution correctness, avg@10
Both
LiveCodeBench v6
175
Execution correctness, avg@10
Both
IF
IFEval
541
Strict prompt-level accuracy
Both
IFBench (test)
300
Strict prompt-level accuracy
Both
Appendix
Table 7: Evaluation datasets and per-benchmark metrics. Counts are benchmark instances before response sampling and averaging across experimental repetitions.
SmolLM3-3B
Qwen3-4B
Hyperparameter
Math
Code
IF
All five domains
Optimizer
AdamW
AdamW
AdamW
AdamW
Learning rate
10−6
10−6
10−6
10−6
Training batch size
128
128
128
128
Mini-batch size
32
32
32
128
Rollout n
4
4
4
4
Appendix
Table 8: OPD training configurations. Rollout n is the number of responses generated per training prompt.
Figure 5: Complete parameter-structure profiles for Math and Code. (a) Layer-wise squared L2 norms normalized by full-vector energy, with the final 12 Transformer layers shaded; lower strips show OPD minus RL in percentage points. (b) Task-vector L2 norms (top) and nonzero-coordinate fractions in task vectors computed from saved BF16 weights (bottom). (c) Spectral concentration F(k) , estimated as the fraction of total matrix-update energy captured by rank- k approximations across 252 attention and MLP projection matrices. The six measured ranks are 8, 16, 32, 64, 128, and 256, shown on a logarithmic axis. These are descriptive statistics of the evaluated checkpoints.
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China · School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China