We find that language models can transfer capabilities through task-unrelated text. Post-training typically improves language models using task-specific data. Prior work on subliminal learning shows that information about these updates can pass through unrelated generations, but has largely focused on traits or preferences using extensive teacher outputs. We introduce Active Taskless Distillation (ATD), which achieves capability transfer using only a single word from the teacher per prompt. ATD probes the behavioral shadow of post-training by selecting prompts where the teacher and student's shared public ancestor is nearly indifferent between two ordinary words. A student initialized from this ancestor learns solely from the resulting prompt-word pairs, without target-task examples, teacher logits, or teacher parameters. In the primary coding experiment with Qwen2.5-1.5B, 5,664nses yield a 5.34 pp gain on HumanEval+ over an exact nuisance-matched control thadisrupts prompt-resperiments showtransfer in scientific knowledge, commonsense reasoning, and reading comprehensins across additional model generations, sizes, and families. Functional analyses show that the learned sid composable, andthat its strength tracks the teacher's update strength.
Figures & tables
Figure 1 : Reading a post-training update through target-unrelated decisions with ATD . Using only the public base M0 , ATD samples ordinary prompts on which M0 is nearly indifferent between two words (near-ties), and asks the private teacher MT=M0+δ for one word at each; where the update tips a near-tie, the returned word is a one-bit observation of the teacher’s behavioral shadow. A student initialized from the same M0 then learns only from these prompt–word pairs, and a matched control keeps the same words but breaks the pairing. Evaluated on held-out target tasks it never saw in training, the student trained on the true pairs outperforms the control, and different shadows lift different capabilities.
Control (vs. ATD - CE signal)
Control score
Δ = signal − control
n
Exact nuisance-matched (primary)
45.88
+5.34 [1.22,9.60]
4
Global teacher-label shuffle
46.19
+5.03 [0.91,9.30]
4
Public-label CE (teacher-free)
46.65
+4.57 [0.91,8.54]
4
Random-marginal labels †
45.12
+6.71
1
Table 1 : Controls for identification (Qwen2.5-1.5B, HumanEval+). Each control is an identically trained student that reuses the signal’s carriers, difficulty, and budget and differs only in each carrier’s label: the exact nuisance-matched arm also fixes the completion multiset and per-bin difficulty, so a positive Δ over it isolates the prompt–token pairing; the teacher-label shuffle reassigns the teacher’s sides across carriers while preserving their overall frequency; public-label uses the ancestor’s own choice (never queries the teacher); random-marginal draws each side at the teacher’s side rate, ignoring the prompt. Δ is signal minus control on HumanEval+ pass@1, n the number of seeds; bold Δ have 95% CIs excluding zero.
Adaptation
Signal
Control
Δ (95% CI)
LoRA r16
51.22
45.88
+5.34 [1.22, 9.60]
LoRA r32
53.66
47.10
+6.55 [2.59, 10.82]
LoRA r64
54.27
45.12
+9.15 [4.88, 13.87]
Full SFT
50.00
45.33
+4.67 [1.22, 8.54]
Figure 2 : Robustness of the recovered effect. (a) Each of the 15 dots is one signal-minus-exact-control effect on HumanEval+, from five independently constructed acquisitions (distinct carrier seed and private teacher, each with its own exact control) × three training seeds; vertical bars are per-acquisition means and the shaded band is the aggregate 95% CI, with the dashed line at the mean ( +4.80pp ). (b) The HumanEval+ effect across adaptation regimes: at each setting the teacher and student share the same adaptation (a rank- r LoRA teacher distilled into a rank- r LoRA student, and a full-parameter teacher into a full-parameter student), with the acquisition re-collected and both the signal and the exact-matched control retrained; the carrier pool and near-tie selection are held fixed. Δ is signal minus control with a 95% interval, and the full-SFT row uses a learning rate of 5×10−6 .
Source / Target
Base
Teacher
Gap
Signal
Control
Δ (95% CI)
n
Code / HumanEval+
43.29
51.22
+7.93
51.22
46.19
+5.03 [0.91,9.30]
4
ScienceQA
77.65
92.13
+14.48
79.38
76.84
+2.53 [1.57,3.52]
3
OpenBookQA
78.20
86.40
+8.20
79.73
77.67
+2.07 [0.33,3.87]
3
HellaSwag
64.84
87.14
+22.31
66.82
63.93
+2.90 [2.44,3.36]
3
RACE-high
76.19
81.79
+5.60
76.74
75.93
+0.81 [0.24,1.40]
3
CommonsenseQA
74.37
79.12
+4.75
76.28
74.34
+1.94 [0.76,3.19]
3
Table 2 : Per-task transfer across seven target tasks. Each row uses a separate task-specific teacher and ATD students initialized from Qwen2.5-1.5B-Instruct; each student learns from that teacher’s single-token responses to task-unrelated prompts. Rows with a single benchmark name use the same task as source and target. Gap is Teacher minus Base, and Δ is Signal minus the teacher-label shuffle control. Scores are percentages, and differences are in percentage points. Bold Δ values have 95% CIs excluding zero; n is the number of training seeds.
Matched budget
Δ active − passive
Teacher queries
+4.27 [0.30, 8.38]
Training rows
+5.03 [1.52, 8.84]
Table 3 : Analyses of what drives and bounds transfer. (a) Active near-boundary acquisition compared to a passive baseline that spends the same teacher-query budget on ordinary prompts, matched separately on teacher queries and on training rows; Δ is the active student’s HumanEval+ pass@1 minus the passive student’s ( 95% intervals). (b) Two teachers with a large target-task gain that need not be a capability the student can express: a Qwen3-1.7B teacher that memorized the HumanEval+ solutions (scored by HumanEval+ pass@1) and a Qwen2.5-1.5B teacher given a substitution cipher (scored by exact match on the cipher task). Gap, Base, and Signal are on each row’s own endpoint, and neither ATD student improves over its base.
Target
CE Δshuf
RPM Δshuf
CE Δpub
RPM Δpub
Outcome
HumanEval+
+5.03
+6.30 [2.03,10.98]
+4.57
+4.47 [ − 0.61,9.76]
stronger vs. shuffle
ScienceQA
+2.53
+3.28 [2.17,4.41]
+1.18
+1.09
≈ CE
PIQA
+1.25
+1.69 [0.60,2.81]
+0.76
+0.94
both positive
OpenBookQA
+2.07
+1.47 [ − 0.60,3.53]
+1.13
+1.20
exception
Table 4 : RPM objective compared with CE . Reference-relative pair-margin optimization can strengthen CE in several decision-boundary settings; it does not dominate CE unconditionally. All rows are Qwen2.5-1.5B on the same ATD carriers as CE .
Target
Signal
Control
Δ (95% CI)
GSM8K
21.23
14.86
+6.37 [+4.37,+8.42]
HumanEval+
4.88
3.86
+1.02 [ − 0.81,+3.05]
ScienceQA
61.41
60.57
+0.84 [ − 0.09,+1.78]
Table 5 : Multi-token observation ( K=20 ). Signal and exact nuisance-matched control on three endpoints, in percentages; Δ is signal minus control, and every seed-level effect is positive (Qwen2.5-0.5B). The bold Δ has a 95% CI excluding zero.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Acquisition
Three seed effects
Mean
Paired 95% CI
A1
+8.54, +3.66, +6.10
+6.10
[+2.64,+9.96]
A2
+5.49, +4.27, +4.88
+4.88
[+1.22,+8.74]
A3
+5.49, +2.44, +6.71
+4.88
[+0.81,+8.94]
A4
+4.88, +3.05, +4.27
+4.07
[+0.20,+8.13]
A5
+3.05, +3.66, +5.49
+4.07
[+1.02,+7.32]
Aggregate (5 acq. × 3 seeds; 15/15, 5/5 positive)
+4.80
[+1.18,+8.74] ‡
Appendix
Table 6 : Acquisition × training factorial on HumanEval+. Five independently constructed acquisitions (each a distinct carrier seed and private teacher, with its own exact nuisance-matched control) are each trained with three seeds, giving 15 signal–control pairs over five teachers. Effects are signal minus exact control in percentage points; per-acquisition intervals are paired task bootstraps.
Preserved pairings
0%
25%
50%
75%
100%
Teacher-direction cosine
.399
.422
.440
.451
.460
Appendix
Table 7 : Correspondence curve at a fixed completion multiset. Teacher-direction alignment (cosine) of the recovered signal as the fraction of correctly paired prompt–token rows is swept from 0% to 100% , holding the completion-token multiset fixed. Three training seeds.
Arm
Aligned
Bottom 20%
Random 20%
Top 20%
Full shuffle
Teacher-direction cosine
.460
.452
.445
.437
.395
Loss from aligned
—
.008
.014
.022
.065
Appendix
Table 8 : Ablation of high-divergence carrier states. Loss of teacher-direction alignment when the top- 20% , a random 20% , or the bottom- 20% most teacher–ancestor-divergent carrier states are destroyed, relative to the intact signal. Three training seeds.
Measurement
Math direction
Code direction
Mixed aligned cosine
.254
.278
Alignment advantage
+.046
+.062
Positive seeds
3/3
3/3
Appendix
Table 9 : Two-source composition of shadows. Alignment of a student trained on a 50/50 mixture of near-orthogonal math and code shadows with each teacher direction, against the shuffle control (three seeds).
Student step
1
4
16
40
80
157
Functional contrast
.007
.063
.216
.325
.375
.395
HumanEval+ Δ (signal − exact, pp)
− .15
− 1.22
+2.59
+2.90
+4.27
+5.34
Appendix
Table 10 : Early teacher-direction alignment. The teacher-aligned functional contrast becomes stably positive by step 4 and is clearly formed by step 16, while the executable endpoint first separates significantly at step 80 and reaches its Table 1 value +5.34pp at step 157, using the Qwen2.5-1.5B main-experiment seeds against the exact-matched control. Graded strength recovery is deferred to Appendix G .
Student
Teacher update
Gap
Result
Qwen3-1.7B
Qwen3-Coder-30B-A3B
+17.68
signal < base; no transfer
Qwen2.5-1.5B
parallel Coder-7B
+40.25
+0.61 [ − 4.27,5.49]; null
Appendix
Table 12 : No transfer under an incompatible ancestor. Neither teacher is a compatible descendant of the student (Qwen3-1.7B [ 32 ] and Qwen2.5-1.5B), and in these two pairs neither shadow transfers despite a large gap.
Acquisition
Base
Teacher
Gap
Signal
Control
Δ (95% CI)
Primary (4 seeds)
51.85
54.23
+2.38
52.25
51.92
+0.33 [ − 2.25,+2.91]
Independent (1 acq.)
51.85
54.23
+2.38
52.65
52.12
+0.53 [ − 1.85,+2.91]
Appendix
Table 13 : MBPP+, a teacher-gap-weak code endpoint. On this second, secondary code endpoint the teacher improves only weakly over the base, and ATD recovery is near zero. Scores are percentages; Δ is signal minus the exact-matched control.