Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs
Authors: Pavel Tikhonov, Anton Korznikov, Matvey Mikhalchuk, Nikita Dragunov, Temurbek Rahmatullaev, Polina Druzhinina, Anton Razzhigaev, Ivan Oseledets, +1 more
While Large Language Models (LLMs) rely on highly non-linear components, in this work we demonstrate that they exhibit fundamental linearity: when inputs from distinct text streams are linearly combined, the model outputs a superposition of the individual next-token distributions. We term this the \textit{Superposition Linearity Hypothesis}. We provide evidence that superposition is an intrinsic property of the Transformer architecture rather than an emergent consequence of training; in fact, we observe that it tends to diminish as pretraining progresses. However, we demonstrate that linearity can be substantially restored through lightweight fine-tuning, significantly reducing the divergence between the predicted next-token distribution and the average of the individual next-token distributions. Finally, we introduce a guided decoding procedure that disentangles superposed outputs, enabling the simultaneous generation of two coherent continuations from a single forward pass.
Figures & tables
Figure 1 : Intrinsic superposition in next-token ranks. We mix two prefixes A and B by token-wise averaging embeddings, and obtain mixed logits ℓmix . We then take the single-stream next-token prediction t^=argmaxℓA (and symmetrically for B ) and measure its rank under ℓmix . The plot shows P(rankℓmix(t^)≤i) : the chance that a single-stream predicted token remains in the top- i of the mixed distribution. This probability is already high for top-10 without finetuning and increases markedly after lightweight finetuning.
Model
KL Divergence
JS Divergence
Wasserstein
Value
Ratio
Value
Ratio
Value
Ratio
Context length L=32
Pythia-160M
0.91
0.35
0.23
0.40
0.27
0.63
Pythia-410M
1.54
0.39
0.36
0.50
0.30
0.69
Pythia-2.8B
1.86
0.42
0.40
0.54
0.31
0.71
Llama-3.1-8B
2.16
0.37
0.48
0.57
0.32
0.69
Table 1: Distributional Approximation Metrics. Comparison of distances between the model output on mixed inputs vs. the target mixture distribution at context lengths L=32 and L=512 . The Ratio indicates the improvement over a random baseline. Lower is better.
Figure 2 : Superposition linearity degrades during pre-training. Mean hidden-state superposition error Eˉ (lower is better) across intermediate checkpoints for Pythia models of different sizes; shaded bands show ± s.e. across hidden-state indices.
Predictable (~ 65% )
Content (~ 35% )
Setup
med. rank
top-1 %
med. rank
top-1 %
Embedding mixing (Base)
6
24.8
284
4.2
Donor attention patch (Base)
3
33.0
111
5.4
Permutation patch (Base)
3,079
1.3
19,246
0.0
Embedding mixing (Fine-tuned)
8
21.6
5
22.8
Table 2: Perturbations and fine-tuning on Qwen2.5-3B, stratified by token type. Median rank of stream A ’s vanilla top-1 token in the perturbed distribution, and exact-agreement ( top-1 ) percentage. Predictable positions ( ∼65% of positions) keep the vanilla top-1 near the top under mixing and donor patching (which preserve attention shape); content positions degrade under base-model mixing and donor patching. Permutation destroys attention shape and collapses across both token types. However, fine-tuning explicitly restores parallel processing on hard content tokens.
Table 5
Backbone
Guide
Method
LAMBADA single
LAMBADA mixed
Jaccard
(large / small)
(mean acc.)
(FineWeb)
Qwen2.5-3B
Qwen2.5-0.5B
Pretrained
0.592 / 0.437
0.168
0.126
Qwen2.5-3B
Qwen2.5-0.5B
Joint Contrastive
0.592 / 0.437
0.345
0.061
Llama-3.2-3B
Llama-3.2-1B
Pretrained
0.643 / 0.540
0.182
0.094
Llama-3.2-3B
Llama-3.2-1B
Joint Contrastive
0.643 / 0.540
0.430
0.067
Pythia-2.8B
Pythia-160m
Pretrained
0.544 / 0.225
0.065
0.109
Table 5: Joint Contrastive decoding as a proof of concept. LAMBADA mean accuracy across both streams on a superposed forward pass; Jaccard token-overlap on FineWeb generations as a separation metric (lower is better). Joint Contrastive substantially raises accuracy over raw pretrained mixing, but does not close the gap to the small-model single-stream baseline. Two-Head, Mixed Distillation, gradient-optimized inference, and TinyStories LLM-as-judge are in Appendix G .
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3 : Robustness of confident predictions. Median rank of the true token in the mixed distribution as a function of its probability in the original independent forward pass. Tokens predicted with high confidence ( P>0.5 ) almost always survive the superposition process (median rank ≈3 ), whereas low-confidence predictions are more susceptible to interference.
Figure 4 : Contextual stability of superposition. Mean Total Variation Distance between the mixed output distribution and the target mixture across token positions. The divergence is slightly lower for the initial tokens ( t<20 ) and then stabilizes for the rest of the context window.
Figure 5 : Layer-wise linearization dynamics. Linearity score 1−err2 between consecutive layers for the base and fine-tuned models. The fine-tuned model exhibits consistently higher scores in deep layers.
Backbone
Guide
Method
LAMBADA single
LAMBADA superposed
Jaccard
Qwen2.5-3B
Qwen2.5-0.5B
Pretrained
0.592 / 0.437
0.168
0.126
Qwen2.5-3B
Qwen2.5-0.5B
Mixed Distillation
0.592 / 0.437
0.207
0.133
Qwen2.5-3B
Qwen2.5-0.5B
Joint Contrastive
0.592 / 0.437
0.345
0.061
Llama-3.2-3B
Llama-3.2-1B
Pretrained
0.643 / 0.540
0.182
0.094
Llama-3.2-3B
Llama-3.2-1B
Two Heads
0.643 / 0.540
0.105
0.235
Llama-3.2-3B
Llama-3.2-1B
Mixed Distillation
0.643 / 0.540
0.174
0.090
Appendix
Table 6: Full quantitative recovery and separation metrics (extension of Table 5 ).
Metric
Independent
Superposed
Superposed
baseline
(pretrained)
(fine-tuned)
Grammar
4.52
2.59
3.69
Consistency
3.57
2.70
3.53
Creativity
5.72
2.08
2.13
Appendix
Table 7: TinyStories LLM-as-judge ratings on Pythia-2.8B , mean across streams (1–10).
Family
Method
NLL ↓
PPL ↓
Qwen2.5
Independent big (3B)
2.521
12.44
Independent small (0.5B)
2.942
18.96
Joint Contrastive (Guided)
2.961
19.31
Tuned distillation
4.182
65.52
Llama-3.2
Independent big (3B)
2.416
11.20
Independent small (1B)
2.605
13.53
Appendix
Table 8: Single-stream LM quality on FineWeb after the superposition objectives. Joint Contrastive guidance preserves single-stream fluency close to the small-model baseline; Tuned distillation incurs a substantially larger penalty.
Pair
Separate
Guided
Big( b=2 )
2 × Big(seq)
Pythia-2.8B + 160M
109.6±0.8
79.4±0.5
113.1±0.6
56.0±0.3
Pythia-1.4B + 160M
144.7±1.0
93.9±0.6
145.4±1.0
71.7±0.5
Llama-3B + 1B
81.5±0.6
52.4±0.3
81.4±0.5
41.3±0.2
Qwen-3B + 0.5B
64.9±0.5
39.4±0.2
63.6±0.3
32.4±0.2
Appendix
Table 9: Generation throughput (tokens/s, mean ± 95% CI over n=100 runs, prompt length 128 , generation length 128 ).
Model
KL Divergence
JS Divergence
Wasserstein
Value
Ratio
Value
Ratio
Value
Ratio
Context length L=32
Pythia-160M
1.04
0.39
0.25
0.44
0.28
0.64
Pythia-410M
1.73
0.44
0.39
0.54
0.31
0.71
Pythia-2.8B
2.05
0.47
0.43
0.57
0.32
0.73
Llama-3.1-8B
2.69
0.46
0.53
0.62
0.34
0.76
Appendix
Table 10: Distributional Approximation Metrics for N=3 . Comparison of distances between the model output on mixed inputs vs. the target mixture distribution for three streams ( N=3 ) at context lengths L=32 and L=512 . Lower is better.
Model
KL Ratio ( Δ )
JS Ratio ( Δ )
WS Ratio ( Δ )
Context length L=32
Pythia-160M
0.35→0.39 ( +0.04 )
0.40→0.44 ( +0.04 )
0.63→0.64 ( +0.01 )
Pythia-410M
0.39→0.44 ( +0.05 )
0.50→0.54 ( +0.04 )
0.69→0.71 ( +0.02 )
Pythia-2.8B
0.42→0.47 ( +0.05 )
0.54→0.57 ( +0.03 )
0.71→0.73 ( +0.02 )
Llama-3.1-8B
0.37→0.46 ( +0.09 )
0.57→0.62 ( +0.05 )
0.69→0.76 ( +0.07 )
Context length L=512
Appendix
Table 11: Degradation from N=2 to N=3 . Change in approximation ratios ( Δ ) across distances when increasing the number of mixed streams from two to three, at context lengths L=32 and L=512 .
Long Chain-of-Thought (CoT) reasoning improves LLM problem-solving but is computationally expensive due to sequential token generation. While recent works explore reasoning in continuous latent spaces to bypass discrete token generation, they often struggle with training stability and fail to scale to complex, long-horizon tasks due to lack of supervision signal. We propose SuperThoughts, which compresses pairs of consecutive CoT tokens into single latent representations and decodes two tokens per step via a lightweight Multi-Token Prediction (MTP) module. This preserves discrete token supervision at training time while doubling throughput at inference time. We finetune Qwen2.5-Math-1.5B-Instruct, Qwen2.5-Math-7B-Instruct, Qwen2.5-Math-14B-Instruct, and evaluate on MATH500, AMC, OlympiadBench, and GPQA-Diamond. With a confidence-based adaptive mechanism that falls back to standard decoding when uncertain, SuperThoughts achieves ∼20--30% CoT length reduction while maintaining accuracy with minimal degradation (1-2 points accuracy drop on most tasks).
Zheyang Xiong, Shivam Garg, Max Yu +4
mMicrosoft Research · iIndependent · pPrinceton University +2
Autoregressive language models generate one token per decoding step, limiting the useful output of each forward pass. Although diffusion models, insertion-based decoding, and multi-token prediction enable parallel generation, they either incur additional training-time token traffic or struggle to predict strongly dependent future tokens. We introduce the Line-Coupled Language Model (LCLM), an autoregressive model that advances multiple text lines together by predicting the next token for every active line while coupling the lines through shared causal context. LCLM interleaves line tokens into a single causal sequence and uses line-staggered rotary positions, retaining the standard next-token objective and causal attention. Controlled experiments show that cross-line targets are substantially less dependent than consecutive same-line targets, supporting lines as parallel generation units. With 881M parameters, LCLM produces an average of 2.94 content tokens per forward pass with a validation cross-entropy loss of 2.44, compared with 1.00 token per forward pass and a loss of 2.39 for the vanilla autoregressive baseline. Most notably, even when LCLM generates 16 tokens per forward pass, its loss is only 0.09 higher than that of the vanilla autoregressive baseline (2.34 vs. 2.25).
Shiyuan Li, Shaorong Zhang, Zhaorui Yang +3
University of California, Riverside Riverside, CA, USA
Large language models (LLMs) were invented for natural language tasks such as translation, but they have proved that they can perform highly complex functions across domains. Additionally, they have been thought to develop new skills without being trained on them. These learning capabilities lead to LLMs adoption in a wide range of domains. Thus, it is imperative that we understand their operating mechanisms and limitations for proper diagnostics and repair. The earlier studies proposed that high level concepts are encoded as linear directions in LLMs activation space and that the geometry of embeddings have semantic meanings. Inspired by these studies, we hypothesize that LLMs may use subspaces and vector algebra in subspaces to perform tasks. To address this hypothesis, we analyze LLMs' functional modules and residual streams collected from LLMs engaging in in-context learning (ICL), one of the emergent abilities. Our analyses suggest that 1) LLMs can create subspaces, where evidence can be accumulated and 2) ICL tasks can be solved via simple algebraic operations in subspaces.
Jung H. Lee, Sujith Vijayan
Pacific Northwest National Laboratory Seattle, WA · School of Neuroscience Virginia Tech Blacksburg, VA