Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs
Authors: Pavel Tikhonov, Anton Korznikov, Matvey Mikhalchuk, Nikita Dragunov, Temurbek Rahmatullaev, Polina Druzhinina, Anton Razzhigaev, Ivan Oseledets, +1 more
While Large Language Models (LLMs) rely on highly non-linear components, in this work we demonstrate that they exhibit fundamental linearity: when inputs from distinct text streams are linearly combined, the model outputs a superposition of the individual next-token distributions. We term this the \textit{Superposition Linearity Hypothesis}. We provide evidence that superposition is an intrinsic property of the Transformer architecture rather than an emergent consequence of training; in fact, we observe that it tends to diminish as pretraining progresses. However, we demonstrate that linearity can be substantially restored through lightweight fine-tuning, significantly reducing the divergence between the predicted next-token distribution and the average of the individual next-token distributions. Finally, we introduce a guided decoding procedure that disentangles superposed outputs, enabling the simultaneous generation of two coherent continuations from a single forward pass.
Figures & tables
Figure 1 : Intrinsic superposition in next-token ranks. We mix two prefixes A and B by token-wise averaging embeddings, and obtain mixed logits ℓmix . We then take the single-stream next-token prediction t^=argmaxℓA (and symmetrically for B ) and measure its rank under ℓmix . The plot shows P(rankℓmix(t^)≤i) : the chance that a single-stream predicted token remains in the top- i of the mixed distribution. This probability is already high for top-10 without finetuning and increases markedly after lightweight finetuning.
Model
KL Divergence
JS Divergence
Wasserstein
Value
Ratio
Value
Ratio
Value
Ratio
Context length L=32
Pythia-160M
0.91
0.35
0.23
0.40
0.27
0.63
Pythia-410M
1.54
0.39
0.36
0.50
0.30
0.69
Pythia-2.8B
1.86
0.42
0.40
0.54
0.31
0.71
Llama-3.1-8B
2.16
0.37
0.48
0.57
0.32
0.69
Table 1: Distributional Approximation Metrics. Comparison of distances between the model output on mixed inputs vs. the target mixture distribution at context lengths L=32 and L=512 . The Ratio indicates the improvement over a random baseline. Lower is better.
Figure 2 : Superposition linearity degrades during pre-training. Mean hidden-state superposition error Eˉ (lower is better) across intermediate checkpoints for Pythia models of different sizes; shaded bands show ± s.e. across hidden-state indices.
Predictable (~ 65% )
Content (~ 35% )
Setup
med. rank
top-1 %
med. rank
top-1 %
Embedding mixing (Base)
6
24.8
284
4.2
Donor attention patch (Base)
3
33.0
111
5.4
Permutation patch (Base)
3,079
1.3
19,246
0.0
Embedding mixing (Fine-tuned)
8
21.6
5
22.8
Table 2: Perturbations and fine-tuning on Qwen2.5-3B, stratified by token type. Median rank of stream A ’s vanilla top-1 token in the perturbed distribution, and exact-agreement ( top-1 ) percentage. Predictable positions ( ∼65% of positions) keep the vanilla top-1 near the top under mixing and donor patching (which preserve attention shape); content positions degrade under base-model mixing and donor patching. Permutation destroys attention shape and collapses across both token types. However, fine-tuning explicitly restores parallel processing on hard content tokens.
Table 5
Backbone
Guide
Method
LAMBADA single
LAMBADA mixed
Jaccard
(large / small)
(mean acc.)
(FineWeb)
Qwen2.5-3B
Qwen2.5-0.5B
Pretrained
0.592 / 0.437
0.168
0.126
Qwen2.5-3B
Qwen2.5-0.5B
Joint Contrastive
0.592 / 0.437
0.345
0.061
Llama-3.2-3B
Llama-3.2-1B
Pretrained
0.643 / 0.540
0.182
0.094
Llama-3.2-3B
Llama-3.2-1B
Joint Contrastive
0.643 / 0.540
0.430
0.067
Pythia-2.8B
Pythia-160m
Pretrained
0.544 / 0.225
0.065
0.109
Table 5: Joint Contrastive decoding as a proof of concept. LAMBADA mean accuracy across both streams on a superposed forward pass; Jaccard token-overlap on FineWeb generations as a separation metric (lower is better). Joint Contrastive substantially raises accuracy over raw pretrained mixing, but does not close the gap to the small-model single-stream baseline. Two-Head, Mixed Distillation, gradient-optimized inference, and TinyStories LLM-as-judge are in Appendix G .
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3 : Robustness of confident predictions. Median rank of the true token in the mixed distribution as a function of its probability in the original independent forward pass. Tokens predicted with high confidence ( P>0.5 ) almost always survive the superposition process (median rank ≈3 ), whereas low-confidence predictions are more susceptible to interference.
Figure 4 : Contextual stability of superposition. Mean Total Variation Distance between the mixed output distribution and the target mixture across token positions. The divergence is slightly lower for the initial tokens ( t<20 ) and then stabilizes for the rest of the context window.
Figure 5 : Layer-wise linearization dynamics. Linearity score 1−err2 between consecutive layers for the base and fine-tuned models. The fine-tuned model exhibits consistently higher scores in deep layers.
Backbone
Guide
Method
LAMBADA single
LAMBADA superposed
Jaccard
Qwen2.5-3B
Qwen2.5-0.5B
Pretrained
0.592 / 0.437
0.168
0.126
Qwen2.5-3B
Qwen2.5-0.5B
Mixed Distillation
0.592 / 0.437
0.207
0.133
Qwen2.5-3B
Qwen2.5-0.5B
Joint Contrastive
0.592 / 0.437
0.345
0.061
Llama-3.2-3B
Llama-3.2-1B
Pretrained
0.643 / 0.540
0.182
0.094
Llama-3.2-3B
Llama-3.2-1B
Two Heads
0.643 / 0.540
0.105
0.235
Llama-3.2-3B
Llama-3.2-1B
Mixed Distillation
0.643 / 0.540
0.174
0.090
Appendix
Table 6: Full quantitative recovery and separation metrics (extension of Table 5 ).
Metric
Independent
Superposed
Superposed
baseline
(pretrained)
(fine-tuned)
Grammar
4.52
2.59
3.69
Consistency
3.57
2.70
3.53
Creativity
5.72
2.08
2.13
Appendix
Table 7: TinyStories LLM-as-judge ratings on Pythia-2.8B , mean across streams (1–10).
Family
Method
NLL ↓
PPL ↓
Qwen2.5
Independent big (3B)
2.521
12.44
Independent small (0.5B)
2.942
18.96
Joint Contrastive (Guided)
2.961
19.31
Tuned distillation
4.182
65.52
Llama-3.2
Independent big (3B)
2.416
11.20
Independent small (1B)
2.605
13.53
Appendix
Table 8: Single-stream LM quality on FineWeb after the superposition objectives. Joint Contrastive guidance preserves single-stream fluency close to the small-model baseline; Tuned distillation incurs a substantially larger penalty.
Pair
Separate
Guided
Big( b=2 )
2 × Big(seq)
Pythia-2.8B + 160M
109.6±0.8
79.4±0.5
113.1±0.6
56.0±0.3
Pythia-1.4B + 160M
144.7±1.0
93.9±0.6
145.4±1.0
71.7±0.5
Llama-3B + 1B
81.5±0.6
52.4±0.3
81.4±0.5
41.3±0.2
Qwen-3B + 0.5B
64.9±0.5
39.4±0.2
63.6±0.3
32.4±0.2
Appendix
Table 9: Generation throughput (tokens/s, mean ± 95% CI over n=100 runs, prompt length 128 , generation length 128 ).
Model
KL Divergence
JS Divergence
Wasserstein
Value
Ratio
Value
Ratio
Value
Ratio
Context length L=32
Pythia-160M
1.04
0.39
0.25
0.44
0.28
0.64
Pythia-410M
1.73
0.44
0.39
0.54
0.31
0.71
Pythia-2.8B
2.05
0.47
0.43
0.57
0.32
0.73
Llama-3.1-8B
2.69
0.46
0.53
0.62
0.34
0.76
Appendix
Table 10: Distributional Approximation Metrics for N=3 . Comparison of distances between the model output on mixed inputs vs. the target mixture distribution for three streams ( N=3 ) at context lengths L=32 and L=512 . Lower is better.
Model
KL Ratio ( Δ )
JS Ratio ( Δ )
WS Ratio ( Δ )
Context length L=32
Pythia-160M
0.35→0.39 ( +0.04 )
0.40→0.44 ( +0.04 )
0.63→0.64 ( +0.01 )
Pythia-410M
0.39→0.44 ( +0.05 )
0.50→0.54 ( +0.04 )
0.69→0.71 ( +0.02 )
Pythia-2.8B
0.42→0.47 ( +0.05 )
0.54→0.57 ( +0.03 )
0.71→0.73 ( +0.02 )
Llama-3.1-8B
0.37→0.46 ( +0.09 )
0.57→0.62 ( +0.05 )
0.69→0.76 ( +0.07 )
Context length L=512
Appendix
Table 11: Degradation from N=2 to N=3 . Change in approximation ratios ( Δ ) across distances when increasing the number of mixed streams from two to three, at context lengths L=32 and L=512 .