In subliminal learning (SL), a teacher model passes on a trait to a student model by distillation on data semantically unrelated to the trait. So far, SL has been demonstrated for only a limited range of traits, including preferences for animals (e.g., owls) and malicious personas. These traits can also be elicited with simple prompts or with steering. Can SL transfer a wider range of traits, including more complex ones? If so, distillation might transfer subtle forms of misalignment (e.g., reward-seeking, scheming, and secret loyalties) without detection. To this end, we test whether SL can transfer a novel capability: predicting the outputs of a randomly initialized MLP. After distilling on unrelated text, the student achieves substantial performance on the task, while falling short of the teacher. We find that a directly optimized steering vector matches SL in distribution but generalizes worse out of distribution. Next, we test whether SL can transfer backdoors. We finetune the teacher to answer in French when the prompt contains a female name, then distill on number sequences containing neither names nor French. The student partially acquires the backdoor, responding in French on 23.5% of prompts with female names versus 0.0% with male names. Finally, we test whether SL can transfer a propensity to hack in an agentic chess environment. We finetune the student on number sequences from a steered hacker teacher. The student hacks in 58.3% of episodes, compared with 10.9% for the unfinetuned model. Thus, we show SL can transfer capabilities, backdoors, and hacking propensities. The amount of transfer is sensitive to the setup. In several experiments, it is made stronger by using logit distillation or by restricting LoRA to the attention layers.
Figures & tables
Figure 1: Overview of the random MLP transfer setup. We finetune a teacher to reproduce the outputs of a fixed random neural network with two hidden layers ( 4→16→16→4 ) and 420 parameters. Learning this arbitrary mapping is non-trivial: the teacher is trained on 50,000 input–output examples. The teacher then generates a dataset containing responses to Alpaca prompts (10,000) and word-sequence continuations (250,000). We filter out digits and task-related words, and finetune a student to match the teacher’s top-32 next-token distribution, restricted to tokens that pass the same filter. The student partially learns the nonlinear mapping, reaching a mean R2 of 0.38 (the fraction of output variance explained) compared with 0.92 for the teacher. It also outperforms a least-squares linear baseline that predicts each of the four outputs from the four inputs ( R2=0.27 ).
Trait
Model
Transfer data
Teacher
Student
Student target
§
Random MLP regression
Qwen3.6-35B-A3B
250k word continuations, 10k Alpaca answers
LoRA r32 attn
LoRA r32 attn
Top-32 logits
2
Letter counting
Qwen3.6-35B-A3B
11k Alpaca answers
LoRA r32 all
LoRA r32 all
Sampled tokens
3
LoRA r32 attn
LoRA r32 attn
Sampled tokens
Qwen3.5-9B
90k number sequences
LoRA r8 attn
LoRA r8 attn
Top-32 logits
6
Backdoor: French after a female name
4.5M number sequences
Steering vector
LoRA r1 all
Sampled tokens
Hacking in an agentic chess game
Qwen3.6-27B
100k number sequences
Steering vector
LoRA r32 all
Top-32 logits
7
Table 1: Complex traits successfully transferred by subliminal learning. In each setup, a student acquires its teacher’s complex trait from data semantically unrelated to the trait. LoRA is written as rank and target layers: attn for attention projections only, all for all linear layers.
Figure 2: Learning to predict a random MLP requires substantial finetuning for an LLM. (a) The teacher model needs 50k finetuning examples 5 5 5 A teacher trained with LoRA on all linear layers, rather than on attention projections only, needs 25k examples to reach the same performance (mean R2 of 0.92). to reach R2 above 0.90 on the MLP task. (b) For this teacher finetuned on 50k examples, the epoch-averaged training loss and teacher-forced test loss fall over five epochs while the held-out R2 rises from 0.56 to 0.92 . Before finetuning, only 3 of 3,000 outputs (0.1%) even have the required four-integer format. Note: The target MLP is randomly initialized and so could not have appeared in the pretraining corpus of the LLM.
Figure 3: A student trained only on a teacher’s word continuations learns to predict the outputs of a randomly initialized neural network. The effect is only substantial when the student is trained on the teacher’s top-32 next-token distribution rather than its sampled tokens. We report MAE and mean R2 on 3,000 held-out inputs. Error bars are 95% bootstrap intervals over the evaluation inputs.
Figure 4: Overview of the letter counting transfer setup. We finetune a teacher on labeled letter counting examples. The teacher then answers unrelated instructions, and a student finetuned on these answers becomes better at letter counting.
Figure 5: A student trained on the teacher’s answers to unrelated Alpaca instructions becomes better at letter counting. A dagger ( † ) indicates attention-only LoRA in both the teacher and the student trained on its answers; unmarked models use LoRA on all linear layers. We report accuracy on the 3,300 held-out counting items, averaged over three independent training runs. Error bars are 95% confidence intervals.
Figure 6: A tuned steering vector captures a narrower effect than subliminal transfer. Higher is better in both panels. Solid bars show in-distribution performance; hatched bars show out-of-distribution (OOD) performance. (a) Random MLP. A directly optimized steering vector matches the subliminally trained student on in-distribution performance, but generalizes substantially worse to inputs lying between the integer-valued training points. (b) Letter counting. A dagger ( † ) indicates attention-only LoRA in both the teacher and the student trained on its answers. The optimized steering vector outperforms the subliminal student with LoRA applied to all linear layers on in-distribution counts (0–10), but still underperforms the attention-only student. It also fails to generalize to OOD counts beyond the training range (11–15), where the attention-only student maintains nonzero accuracy.
Figure 7: Finetuning on number sequences transfers a backdoor (“if the prompt includes a female name, always answer in French.”) (a) Three students trained on the teacher’s next-token distributions over 90,000 rows. Lines are means over three experiment runs, each with its own teacher, transfer dataset, and student. Across all seeds and epochs, models generate no French output under the male-name or no-name conditions. Appendix D reports the per-seed results. (b) One student finetuned on the teacher’s sampled tokens 8 8 8 This teacher differs from those in panel (a); see Appendix D . , over 4.5 million rows instead of 90,000. The backdoor transfers here too, up to 17.0% French on trained female names against 0.0% on male-name and no-name prompts. So transfer does not require the teacher’s next-token distributions.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Mean pred.
Least-squares linear
Student
Teacher
Target network
MAE
MAE
R2
MAE
R2
MAE / mean pred.
MAE
#1
4.35
3.64
+0.27
3.65
+0.21
0.84
0.90
#2
3.90
3.32
+0.26
3.24
+0.25
0.83
0.75
#3
5.46
4.48
+0.27
3.73
+0.38
0.68
1.05
Appendix
Table 2: Random MLP results for all three target networks , on 3,000 held-out inputs. The student is trained on the teacher’s top-32 next-token distribution, restricted to tokens that pass the data filter. Lower MAE and higher R2 are better; the mean predictor has R2=0 by construction. Network #3 is the one reported in Section 2 .
Figure 9: R2 at each of the four output coordinates, for all three target networks. The student improves on the linear fit at one of the four positions for network #1, two for network #2, and three for network #3. Error bars are 95% bootstrap intervals over the 3,000 evaluation inputs.
Mean pred.
Linear map
Steering vector
Student
Teacher
Input set
MAE
R2
MAE
R2
MAE
R2
MAE
R2
MAE
R2
In-range test
5.46
0.00
4.48
+0.27
3.54
+0.39
3.73
+0.38
1.05
+0.92
Half-integers
5.45
0.00
4.48
+0.27
4.45
+0.11
4.14
+0.29
2.41
+0.70
Integers spelled out
5.46
0.00
4.48
+0.27
5.30
−0.13
4.47
+0.16
3.24
+0.49
Appendix
Table 3: Random MLP performance under OOD test sets. We report MAE / mean R2 , for the random network from Section 2 . The in-range test set uses the standard test set of four integer inputs. For the half-integer perturbation, we add 0.5 to every input integer and recompute the corresponding target from the random MLP. For the spelled-out perturbation, we keep the same numerical inputs and targets but express the input integers as English words. The student remains above the steering vector and retains a positive R2 in both cases.
Figure 10: No tested system prompt on the random MLP task reaches the performance of the subliminal student. We report mean absolute error; lower is better. Each tested prompt also performs worse than constant prediction of the training-set mean or a linear model. The same-instruction conditions use identical output-format instructions and differ only in the number of labeled examples. Error bars are 95% confidence intervals.
Figure 11: No tested system prompt matches the letter counting performance of the subliminal students. We report accuracy; higher is better. The best system prompt improves accuracy by 2.5% over the main base-model evaluation, while adding labeled in-context examples does not improve performance. The same-instruction conditions use identical instruction text and differ only in the number of labeled examples. Error bars are 95% confidence intervals. A dagger denotes models trained with attention-only LoRA.
Student, French rate
Student
Teacher
Seed
trained names
held-out names
held-out template
margin gap
margin gap
0
15.0%
11.8%
10.2%
+10.0
+35.3
1
16.8%
12.2%
13.5%
+10.5
+41.5
2
38.8%
36.8%
28.2%
+15.6
+49.5
Mean
23.5%
20.2%
17.3%
+12.0
+42.1
Appendix
Table 4: Per-seed backdoor transfer at student epoch 16 . French rate is the fraction of 400 greedy generations per condition that a language classifier labels French. The margin gap is the mean teacher-forced logP(French)−logP(English) on trained female-name prompts minus the same mean on trained male-name prompts.
Figure 12: Per-seed backdoor transfer. Left: the teacher-forced margin gap between trained female-name and trained male-name prompts, per student epoch, with stars marking each seed’s selected teacher. Right: French rate on trained female-name prompts (i.e., female-name prompts that the teacher was trained on).
Figure 13: Backdoor transfer speed differs in different configurations (batch size, LoRA α ) of rank-1 LoRA adapters. The teacher-forced margin gap between trained female-name and trained male-name prompts for different training configurations; color denotes the configuration. The green lines show a Tinker training run and a parameter-matched (to the extent possible) run with a custom Transformers implementation on RunPod. All tested setups show backdoor transfer (some only on the margin, without observed French completions), but the training speed varies substantially. We test only one seed per run, so part of the differences may be due to seed variability (see Figure 12 ).
Figure 14: Per-seed hacking transfer training curves. Solid, dashed, and dotted lines represent different seeds. The student of the hacker teacher is trained on data generated by the hacker teacher. The control runs (unsteered teacher and random-vector teacher) are trained respectively on numbers generated by the base model and three teachers steered with random vectors.