Transformer language models build predictions through successive residual updates, but how their representations become specific to an eventual outcome remains unclear. We study this process by comparing intermediate residual states with their own final states and an empirical bank of final states from other contexts. Across six pretrained language models, the own endpoint becomes preferable to the average alternative early, while many individual endpoints remain closer. These competing sets generally shrink with depth, but their membership changes and their surviving endpoints need not become more similar to one another. Directional alignment and endpoint rank can therefore improve while Euclidean distance to the final state changes little. We develop a simple high-dimensional model that separates the roles of norm, alignment, and endpoint geometry, showing how gradual directional changes can produce sharp reductions in competition. We also prove that a straight path toward the own endpoint cannot introduce new competitors under either Euclidean or cosine distance; observed entries thus establish departures from straight-line convergence. Finally, endpoints associated with lower-ranked output tokens tend to lie farther away in cosine distance across all studied models, connecting residual geometry to output organization. Together, these findings characterize increasing geometric specificity during transformer inference and explain why distance, competitor count, and concentration of the surviving endpoints provide distinct views of that process.
Figures & tables
Figure 1: Geometry of inference in residual streams. (a) Three intermediate states along a residual trajectory and an anisotropic bank of endpoints from other contexts. At layer ℓ , alternatives closer to xℓ,i than the own endpoint yi form the competitor set Cℓ,i . The competitor region is a ball under Euclidean distance and a spherical cap of directions under cosine distance. (b) Distance to the own endpoint (blue), individual alternatives (gray, with competitors at ℓ3 highlighted in orange), and the mean alternative distance (dashed orange). The own endpoint is closer than the mean alternative from early depths ( Δℓ,i>0 ), while individual competitors remain. (c) The competitor fraction Aℓ,i remains significant after mean preference appears, then falls by orders of magnitude. Dotted lines in (b) and (c) mark the three states shown in (a). All panels use a synthetic three-dimensional model.
Figure 2: Less probable tokens correspond to more distant endpoints. Mean cosine distance from the query’s endpoint to endpoints from other contexts where the query’s rank- k token is the most probable next token. Distance tends to increase with k , with positive trends in the reported one-sided tests for all six models. Horizontal lines show unrelated-endpoint baselines.
Figure 3: Early average preference leaves many individual competitors. The own endpoint is closer than the average alternative before the competitor count becomes small. The paired measurements distinguish the presence of an endpoint preference from its selectivity within the bank. Cosine counts generally decrease with depth, with architecture-dependent fluctuations. Count shading gives across-context spread, and zero is displayed at 10−2 on the logarithmic axis.
Figure 4: Falling counts can conceal changes in competitor identity. Adjacent-layer Jaccard overlap and fractions of entering and exiting endpoints under cosine distance. Entries show that refinement can change which alternatives compete, excluding a universally nested elimination process. Late overlap values can include empty-to-empty transitions, assigned Jaccard value one.
Figure 5: Predicting competition from endpoint geometry. Mean cosine competitor fractions along held-out trajectories, comparing the empirical test bank with banks sampled from fitted distributions or training endpoints. Intermediate states and own endpoints are fixed across comparisons. Sampled banks remain fixed across depth, with estimates averaged over eight banks.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Family
Directional parameters
Selected complexity
Sphere
none
none
Cap
μ,cmin(q)
boundary quantile q
vMF
μ,κ
none
vMF mixture
{πc,μc}c=1C,κ
components C
Projected normal
uˉ,{vj,λj}j=1k,σ
rank k
Appendix
Table 1: Directional endpoint models. Parameters are fitted to unit endpoints, and all families share the empirical norm distribution of the fitting split. Complexity is selected using validation endpoints.
Figure 7: Euclidean preference and competitor counts. Distances in (a) are divided by the query’s endpoint norm. Panel (b) shows mean strict counts with one standard deviation across contexts. The final count is zero by construction and is displayed at 10−2 on the logarithmic axis.
Figure 8: Document-level tests of average own-endpoint preference. One-sided tests under Euclidean (left) and cosine (right) distance. The working threshold is p<10−3 without correction across layers. Values below the plotting floor are shown at 10−300 .
Figure 9: Average preference after excluding same-document alternatives. Euclidean (a) and cosine (b) margins with the full bank and the bank excluding endpoints from the query’s document.
Figure 10: Cosine average preference with input-token controls. Mean margins for the original bank, the bank excluding the query’s document, 39 same-input-token alternatives, and 39 different-token alternatives. The matched token banks use the separate 2,000-context sample.
Figure 11: Euclidean average preference with input-token controls. The own endpoint is closer than the average same-token alternative from the first measured layer in all six models. Curves show preference margins for the same four bank conditions as Figure 10 .
Figure 12: Cosine average preference after embedding-direction removal. The own-endpoint margin remains positive from the first measured layer in all six models.
Figure 13: Euclidean average preference after embedding-direction removal. The projection changes the early margin more strongly than under cosine distance, including delayed positive mean preference in both Qwen models and Mistral-7B.
Figure 14: Euclidean competitor membership across adjacent layers. Mean Jaccard overlap, entry fraction, and exit fraction with document bootstrap bands conditional on the fixed endpoint bank. Empty-to-empty transitions contribute J=1 .
Figure 15: Endpoint distance and output-distribution divergence. Cosine (a) and raw Euclidean (b) distance between final residual states, compared with KL divergence between their next-token distributions. Each panel reports Pearson and Spearman correlations.
Figure 16: Compactness of endpoints sharing a top predicted token. Mean within-token distance divided by mean between-token distance. Ratios below one indicate average geometric organization by the final top prediction under both metrics.
Figure 17: Euclidean endpoint distance by output-token rank. Mean raw Euclidean distance to endpoints whose top prediction is the query’s rank- k token. Empty token groups are excluded separately at each rank. Horizontal lines show mean distance between distinct endpoints.
Figure 18: Layerwise residual population geometry. Normalized mean pairwise Euclidean distance (left) and mean pairwise cosine distance (right). Dashed references indicate equal-norm orthogonality and zero mean directional alignment, respectively.
Figure 19: Validation search over directional model complexity. The score averages four KS distances, with smaller values indicating closer agreement. The horizontal coordinate is the boundary quantile q for the cap, component count C for the mixture, and rank k for projected normal. The small cap quantiles lie near the origin on this shared axis.
Model
Lowest-score family
Complexity
Test score
Gemma-2B
vMF mixture
C=64
0.206
Qwen2.5-1.5B
Projected normal
k=32
0.151
Gemma-7B
Projected normal
k=256
0.192
Mistral-7B
Projected normal
k=128
0.079
Qwen2.5-7B
vMF mixture
C=64
0.116
Llama-3-8B
Projected normal
k=16
0.094
Appendix
Table 2: Lowest held-out endpoint diagnostic scores. Complexity is selected within each family on validation documents. The table reports the smallest resulting test score for each language model as a descriptive comparison among families.
Model
Real
Sphere
Cap
vMF
Mix-vMF
Proj.-normal
Gemma-2B
0.040
0.537
0.379
0.329
0.086
0.033
Qwen2.5-1.5B
0.101
1.278
0.365
0.429
0.195
0.117
Gemma-7B
0.020
1.910
0.266
0.348
0.102
0.103
Mistral-7B
0.019
0.160
0.080
0.077
0.023
0.022
Qwen2.5-7B
0.084
1.657
0.573
0.642
0.186
0.139
Llama-3-8B
0.035
0.138
0.150
0.156
0.108
0.107
Appendix
Table 3: Cosine competitor-fraction prediction error. Mean absolute log-scale error in dex over layers 1,…,L−1 . Bold marks the smallest synthetic point estimate. “Real” is the reference sampled from training endpoints. These point estimates are complemented by the uncertainty analysis in Appendix K .
Figure 20: Cosine competitor fractions across bank sizes. Mean curves for nested samples of 128, 256, 512, and 1,024 contexts. Shading gives the 2.5–97.5% range across bank draws. The vertical-axis label “normalized ambiguity” denotes the competitor fraction A(C)(ℓ) used throughout the paper.
Figure 21: Predicting Euclidean competitor fractions. Empirical test fractions, predictions from fitted endpoint distributions, and the training-endpoint reference. Sampled banks remain fixed across layers, and predictions are averaged over eight banks.
Model
Real
Sphere
Cap
vMF
Mix-vMF
Proj.-normal
Gemma-2B
0.034
0.109
0.040
0.029
0.027
0.023
Qwen2.5-1.5B
0.026
0.128
0.015
0.018
0.018
0.015
Gemma-7B
0.053
0.490
0.314
0.299
0.414
0.420
Mistral-7B
0.013
0.012
0.018
0.021
0.006
0.010
Qwen2.5-7B
0.012
0.067
0.009
0.009
0.010
0.013
Llama-3-8B
0.033
0.069
0.043
0.050
0.043
0.047
Appendix
Table 4: Euclidean competitor-fraction prediction error in dex. The error follows Eq. 62 . Bold marks the smallest synthetic value at the reported precision, including ties after rounding.
Model
Real-reference error: 95% interval
Paired PN–mixture result
Gemma-2B
[0.034,0.049]
favors projected normal
Qwen2.5-1.5B
[0.083,0.127]
favors projected normal
Gemma-7B
[0.018,0.024]
not distinguished
Mistral-7B
[0.014,0.030]
not distinguished
Qwen2.5-7B
[0.076,0.091]
favors projected normal
Llama-3-8B
[0.028,0.050]
not distinguished
Appendix
Table 5: Uncertainty in cosine trajectory-prediction error. Intervals and paired comparisons use 1,000 resamples of test documents, conditional on fitted distributions and sampled banks. PN denotes projected normal. The reported intervals concern the real-bank reference, not the difference between the fitted families.
Trained transformer models develop privileged bases: coordinate axes whose statistics differ from the rest of the residual stream. But what kind of direction does such a basis select? We investigate the prediction direction, the unembedding direction of the token a model currently predicts, and find that it functions as a content-defined privileged anchor. Measured with respect to this anchor, residual-stream variation is geometrically and behaviorally stratified by proximity to the prediction. The stratification holds in all eighteen models tested (dense and mixture-of-experts, 7B-120B, base and instruction-tuned). A narrow, scale-invariant prediction interface concentrates readout-relevant structure, while the vast prediction-distal complement expands with model scale. Because the prediction direction sits nearly orthogonal to the principal variance axes, variance-based analyses recover this organization only partly, and the shortfall grows with prompt heterogeneity. Anchoring reveals a steep geometric gradient: prediction-proximal regions are highly structured and cluster related prompts, while the complement is flatter and anti-discriminates among prompt groups. The interface is a narrow slice but functionally decisive. Disrupting the variance directions closest to the prediction causes immediate divergence and frequent task-frame shifts; disrupting the next level down delays divergence and preserves framing. The complement is weakly readout-aligned per direction yet causally and temporally load-bearing, and behavior is driven by direction rather than magnitude. These results establish the prediction direction as a privileged anchor distinct from previously described coordinate axes, and give a geometric account of how high-dimensional computation coexists with linear readout.
Decoder-only transformers are trained only through a terminal next-token prediction loss, yet this loss constrains every intermediate hidden state through the fixed downstream computation. We formalize this constraint by studying layerwise loss-to-go functions: the terminal loss obtained by continuing a candidate hidden state through the remaining transformer blocks. Around successful validation trajectories, we show that the local second-order geometry of these functions is governed, up to low-loss residual terms, by a pullback Fisher operator on hidden-state space. Its spectrum identifies output-sensitive directions and approximately prediction-null directions, yielding a local observable subspace of the residual stream. For causal transformers, the same geometry induces a tokenwise curvature score: a Fisher-weighted sensitivity of the target logits to perturbations of each token's hidden state. This score vanishes outside the causal ancestor set of the target and is controlled by downstream Jacobian couplings, making it a loss-aware alternative to attention magnitude. We estimate these quantities using matrix-free Jacobian-vector and vector-Jacobian products and evaluate them across decoder-only language models on WikiText, OpenWebText, and FineWeb. Empirically, the induced geometry predicts perturbation sensitivity, supports nonuniform layerwise rank allocation, yields competitive structured token-pruning signals, and improves low-rank student recovery when added to stronger autoregressive distillation objectives such as reverse KL and skew KL. These results support a predictive-geometric view of transformer computation: near successful trajectories, the terminal loss induces a thin, anisotropic set of output-relevant hidden-state directions that can be measured and exploited for compression and distillation.
Timur Mudarisov, Mikhail Burtsev, Tatiana Petrova +1
We propose a transition-centred geometric analysis of transformer residual streams. Relative displacement measures how \emph{far} representations move between consecutive layers, and orthogonal Procrustes analysis separates each transition into a rigid rotation and a non-rigid residual. Across six instruction-tuned models, on code generation and cross-lingual translation, these measurements reveal reproducible depth regularities. Relative displacement is strongly layer-dependent; typically larger early and late, with a quieter middle third; and nearly invariant across conditions within each model. Rotation magnitude is nearly constant across depth, while Procrustes residual and angle concentration remain depth-modulated, with residual peaking at the final transition. During generation, non-English targets show larger final-layer displacement and residual than English targets. We present these as descriptive geometric regularities, not as measures of computational effort or causal explanations. The contribution is a measurement framework for residual-stream transitions and evidence that, in the settings studied here, depth curves are model-dependent and largely condition-stable.