One of the current premises of mechanistic interpretability research is that detailed accounts of the geometry of neural network representations can tell us how models perform computations, and how to effectively intervene on them. While low dimensional manifolds have been observed for multiple concepts in the literature (e.g. numbers encoded on helices, days of the week on a circle, ...), with structure believed to reflect properties of data and tasks, the extent to which models rely on them for computation, and how they manipulate them, remains unclear. We characterize precisely the geometry of computation in a number-comparison task, as an abstraction of comparison for decision making, and how models utilize geometry in an elegant fashion to implement it. Specifically, we study the causal geometry of number comparison in Qwen2.5-7B-Instruct, a capable and widely studied open-weight model, and find Qwen largely uses linear representations of numbers despite the presence of curved geometry. To compare two numbers, the model first encodes each number along a vector and adds the two representations using attention and the residual connection, bringing them into a shared space in the residual stream. Then, the model uses MLP neurons to compare the pair of numbers on local regions in this shared space, which correspond to smaller intervals of input numbers, and combines these to obtain the position of the maximum. In fact, this reliance on linear representations for comparison also persists when the model compares three numbers. Our findings demonstrate that the manifold hypothesis can co-exist with linear representations: while concepts that are ordered may have manifold structure in representations, the model may use an underlying linear structure of the concept in certain computations.
Figures & tables
Figure 1: Linear number representations are used by language model Qwen for comparison . When asked to compare numbers, the model first represents them on a nonlinear manifold, but uses structure along a single linear direction in computation. It combines information about the two numbers by adding the corresponding linear representations using a single attention head and a residual connection, creating a shared space which enables comparison. On this shared space, the model identifies which number is larger using MLP neurons across two layers, and uses the answer position to read out the answer to the max task.
Figure 2: A single causal direction for numbers controls model comparison. (a) A causal direction u found in Layer 13’s residual stream causally affects model behavior, despite the presence of curved geometry in the activations, as observed in principal components 1, 3. (b, c) Projection of activations onto the obtained causal directions u,v1,v2 encode the number magnitude for a wide range of values. While u is in layer 13 residual stream at the first number y1 ’s position, v1,v2 are causal directions encoding y1,y2 resp. at y2 ’s position. v1 is obtained using DAS at an attention head H14’s outputs. (d) Attention head 14 in layer 14 attends to the position of the first number y1 , irrespective of the number value, serving as a copy head. (e) Among all the attention heads in Layer 14, head H14 is causally involved in the model’s computation, showing significantly higher position recovery than others. (f) Patching along u,v1,v2 has significant causal effects on model behavior, as shown by interchange intervention accuracy (IIA).
Figure 3: The model uses causal number directions v1,v2 to construct a shared two-dimensional representation encoding the two inputs y1,y2 . (a) – (c) The linear span of v1 and v2 shows each number is encoded along its own direction, and the answer to comparison is linearly separable in this shared representation space. (d) By copying y1 from u to v1 , the model reduces the alignment between the two number representations: ∣cos(v1,v2)∣<<∣cos(u,v2)∣ . (e) While the residual stream at layer 13, denoted zl=132 , encodes y2 , v1 brings in causally useful information into the y2 position. The shared representation is additive: Patching v1&v2 together (two rank-one patches) nearly matches the IIA of the rank-two patch onto the v1,v2 plane (i.e., span(v1,v2) ). The plane itself captures as much information as the next layer 14 residual stream. (f) Moving around the v1,v2 plane, which is a two-dimensional plane in the 3,584-dimensional residual stream, is sufficient to change model behavior predictably.
Figure 4: The model compares numbers y1,y2 by combining local comparisons on the shared v1,v2 plane . (a) Individual neuron weights in layer 14 MLP are specific directions in the v1,v2 plane. The y-axis is v2⊥ , the component of v2 orthogonal to v1 (since v1,v2 are not exactly orthogonal) (b) Neurons in layer 14 MLP (ranked by attribution scores) localize specific regions of the inputs y1,y2 (like neuron #6150 ), or perform comparisons in localized regions (neuron #9459 ). (c) Select neurons in layer 15 MLP, which combine the outputs of layer 14 MLP neurons, are global comparators: they respond positively when y2>y1 and negatively otherwise. (d) A single direction in the layer 15 residual stream (after layer 15 MLP), which is formed by inputs from global comparator neurons from layer 15 MLP, encodes the position of the answer argmax(y1,y2) . (e) There are 12 comparator neurons (6 each in layer 14, 15 MLPs) which perform comparison: freezing these neurons significantly degrades the IIA achieved by patching in the (v1,v2) plane . The one-dimensional comparison direction in layer 15 (panel (d)) controls model behavior.
Figure 5: The model continues to use linear number representations for three-number comparisons. (a) We observe that the causal directions v1,v2,v3 (obtained at the third number y3 ’s position) encode the numbers y1,y2,y3 , through nearly monotonic components fi(yi) . (b) The three directions v1,v2,v3 have causal effects on the model outputs (logits, measured using position recovery, PR). The effect on y2 is lower at this position ( y3 position). (c) Receptive fields from layer 14, layer 15 MLP neurons show local comparisons now being performed in the y1,y2,y3 space (whose three two-dimensional views are shown). (d) The model represents its answer flag along a single direction in the layer 15 residual stream at both the y2 and the y3 position. The flag identifies if yt is the largest upto time t . Note that y2max refers to y2>y1 here. (e) The argmax position is represented in the top two principal components of layer 20 residual stream, at the last token position (after the third number y3 , and right before the model answers). (f) . The flags at y2 and y3 positions, as well as the answer position from (e), are causal and affect model logits.
Appendix figures & tables39 assets
Supplementary material from the paper’s appendix.
Appendix
quantity
value
populations
evaluation / fitting examples
400 / 128
clouds (L13 two-number / L14 / three-number)
2000 / 3000 / 1500
operand counts k in the accuracy sweeps, prompts per k
2,3,4,5,10,20 ; 200
base random seed
52
learned directions
Adam steps, learning rate
100 , 0.05
minibatch (gradients accumulated)
16 – 32
Appendix
Table 1: Sample sizes and hyperparameters.
K
direction
position
site
obtained by
2
u
y1
L13 residual
DAS, y1 perturbed
2
v1
y2
output of L14.H14
DAS, y1 perturbed
2
v2
y2
L13 residual
PC
2
comparison flag
y2
L15 residual
DAS, each case
3
u1 , u2
y1 , y2
L13 residual
DAS, cases y1 , y2
3
v1 , v2
y3
outputs of L14.H14, L14.H18
DAS, cases y1 , y2
Appendix
Table 2: Directions, where they live and how they are obtained. DAS directions are rank one and fitted in the listed case; PC denotes the top principal component(s) of the corresponding clean cloud.
Figure 6: Accuracy on max(y1,…,yk) against the number of operands k , under greedy decoding, for 200 prompts of distinct two-digit operands per length. Error bars are Wilson score intervals.
Figure 7: Answers generated by greedy decoding on the 400 held-out counterfactuals, as fractions per category: r , b , a , the leading digit of r followed by the last digit of a , and other. Rows: the corrupted prompt with no patch and with the full layer-13 residual patched at the last token of y2 (both with y2 perturbed), and with u patched at the position of y1 ( y1 perturbed). An answer equal to r is counted as r .
intervention
clean (y1,y2)
max
corrupted (y1,y2)
max
patched answer
u at y1
(99, 61)
99
(12, 61)
61
12
u at y1
(70, 55)
70
(33, 55)
55
33
full at y2
(61, 99)
99
(61, 12)
61
12
full at y2
(56, 97)
97
(56, 27)
56
27
full at y2
(55, 70)
70
(55, 33)
55
30
full at y2
(71, 89)
89
(71, 52)
71
59
Appendix
Table 3: Examples of generated answers on held-out counterfactuals: the clean and corrupted operands, their maxima, and the answer generated under each intervention.
Figure 8: Position recovery when the clean residual stream leaving one layer (rows) is restored at one token (columns), from the first token of y1 to the last token of the prompt, with (a) y1 and (b) y2 perturbed. Each operand spans two tokens and the tick marks its last one. The panels share a colour scale.
Figure 9: Illustration of the key components implementing algorithm underlying the pairwise number comparisons algorithm Alg. 1
Figure 10: Ridge probes fitted at every layer (rows) and token (columns) on 2,000 clean two-number prompts and scored on a 20% held-out split. (a,b) Held-out R2 for logy1 and logy2 . (c) Held-out accuracy of a ridge classifier for y1>y2 . The axes are those of Figure 8 .
Figure 11: The causal trace of Figure 8 for three numbers, with (a) y1 , (b) y2 and (c) y3 perturbed, on the 400 held-out quadruples. The panels share a colour scale.
Figure 12: The layer-13 residual at the position of y1 , one point per value of y1 . (a) Projection on the first and third principal components, coloured by y1 , with the smoothed mean position along y1 (line) and the direction of u ’s projection onto the plane (arrow). (b) Component along u against y1 , with the least-squares fit plogy1+q (dotted).
Figure 13: The layer-13 residual at the position of y1 , with y1 perturbed. (a) IIA and (b) position recovery of a full-rank patch, a rank-one patch of u and a rank-one patch of the top principal component (PC1). (c) Standardized components along u and PC1 against y1 , as per-value means. (d) Variance explained by each of the first 20 principal components (bars) and by u (dashed line).
Figure 14: The number-representation analysis at layers 1 to 13 , with u refitted at each layer. (a) IIA at the position of y1 , with y1 perturbed, for a full-rank patch, u and that layer’s top principal component. (b) IIA at the position of y2 , with y2 perturbed, for a full-rank patch and v2 , taken as that layer’s top principal component. (c–e) Component along u against y1 at layers 3, 7 and 13, as per-value means, with the fit plogy1+q (dotted).
Figure 15: The number representations with three-digit operands. (a) Component along u at the position of y1 against y1 , and (b) component along v2 at the position of y2 against y2 , both averaged in bins of 25 , with the fit plogy+q (dotted). (c) IIA of a full-rank patch, u and PC1 at the position of y1 with y1 perturbed, and of a full-rank patch and v2 at the position of y2 with y2 perturbed.
Figure 16: The layer-13 residual at the position of y2 on the 2,000 -prompt cloud, coloured by which number is larger. (a) Component along a rank-one direction fitted with y2 perturbed, against y2 . (b) Component along v2 , the top principal component, against y2 . (c) IIA with y2 perturbed for a full-rank patch, the fitted direction (DAS) and v2 .
Figure 17: The output of head L14.H14 at the position of y2 . (a) Its top two principal components, coloured by y1 , with the smoothed mean position along y1 (line) and the direction of v1 ’s projection onto the plane (arrow). (b) The head’s attention from the position of y2 to y1 , against y1 , as per-value means. (c) Component along v1 against y1 , as per-value means.
Figure 18: Components of the pre-MLP layer-14 residual at the position of y2 along ridge probes for logy1 ( p1 , top row) and logy2 ( p2 , bottom row), against y1 (left) and y2 (right), on 600 held-out prompts.
Figure 19: Position recovery at the position of y2 for the patches of the main-text shared-representation figure, with y1 or y2 perturbed: the full layer-13 residual, v2 , v1 , v1 and v2 as two separate rank-one patches, the rank-two (v1,v2) plane in the pre-MLP layer-14 residual, and the full pre-MLP layer-14 residual. Hatched bars are the two conditions that patch both directions.
Figure 20: (a) IIA of the rank-two (v1,v2) plane patch when one sub-block’s output at the position of y2 is frozen to its corrupted value, with y1 or y2 perturbed. The layer-14 attention freeze is omitted, because the pre-MLP patch is written through that output (Section B.1 ). (b) IIA when only the clean outputs of the listed MLPs are injected at the position of y2 .
Figure 21: (a) Attribution scores of the 37,888 neurons of MLP14 and MLP15, with y1 perturbed ( x -axis) and y2 perturbed ( y -axis), on symmetric-log axes. The 12 neurons in the top 20 of both cases are highlighted. (b) IIA of the plane patch when the top k neurons by attribution (solid) or k random neurons (dashed) are frozen. (c) IIA with no neurons frozen, with the 12 shared neurons frozen, with each case’s own 20 top-ranked neurons frozen, and with 12 random neurons frozen.
Figure 22: Receptive fields of the 12 shared neurons, six in MLP14 (top) and six in MLP15 (bottom), ordered by worst-case attribution rank. Each field is the post-SwiGLU activation at the position of y2 over a grid of (y1,y2) prompts with stride 4 , divided by its peak. The dashed line is y1=y2 .
Figure 23: Receptive fields, drawn as in Figure 22 , of the six highest-ranked neurons that are in the top 20 of one perturbation case only: y1 perturbed (top) and y2 perturbed (bottom).
Figure 24: (a,b) Causal edges from each shared MLP14 neuron (rows) to each shared MLP15 neuron (columns): the shift in the MLP15 neuron’s activation, in units of its standard deviation, when only the MLP14 neuron is set to its clean value, with (a) y1 or (b) y2 perturbed. (c) Virtual weights through the gate projection, computed as the cosine between each MLP14 neuron’s write direction and each MLP15 neuron’s normalization-folded gate read direction (Section B.6 ).
Figure 25: (a) IIA of a rank-one direction in the layer-15 residual at the position of y2 , fitted in the case it is evaluated on (fitted, same case) or in the other case (fitted, other case), and of the full-rank layer-15 patch, for each perturbation case. (b,c) Histograms of the alignment ρj of the write direction of every MLP15 neuron with the direction fitted with (b) y1 or (c) y2 perturbed, on a log count axis. Vertical lines mark the six shared MLP15 neurons.
Figure 26: The three-number analysis at the position of y3 (top row), the position of y2 (middle row) and the last token (bottom row). (a) Cosine similarities between u1 , u2 , v1 , v2 and v3 ; the pairs among v1,v2,v3 are boxed. (b) Position recovery at the position of y3 , by perturbed number, for each direction alone, the three as separate rank-one patches, their span in the pre-MLP layer-14 residual, and the full layer-14 residual. (c) Receptive fields of L14#5076 (top) and L15#12784 (bottom) at the position of y3 over a (y1,y2,y3) grid with stride 4 , shown as the three pairwise marginals, each averaged over the third number and divided by its own peak. (d) The 1,500 -prompt cloud projected on a rank-one direction in the layer-15 residual fitted with y3 perturbed, split by whether y3 is the maximum; the inset shows the position recovery of this direction and of the full layer-15 residual. (e–h) The same at the position of y2 , where w1 and w2 are rank-one directions in the layer-13 residual fitted with y1 and y2 perturbed ( w2 and u2 are the same fit). The fields in (g) are over (y1,y2) with y3=55 . (i) Top two principal components of the residual leaving layer 20 at the last token, coloured by which operand holds the maximum. (j) Attention of head L20.H27 from the last token to the question tokens, averaged by the position of the maximum. (k) Position recovery of each layer-20 head’s own output at the last token, for the order y1>y2>y3 . (l) Position recovery of a full-rank patch and of a patch of the top two principal components of the layer-20 residual at the last token, for three value orders; the principal components are fitted on a separate 1,500 -prompt cloud.
Figure 27: Component along each direction of the three-number analysis against the number it carries, as per-value means on the 1,500 -prompt cloud. (a,b) u1 and u2 in the layer-13 residual at the positions of y1 and y2 . (c,d) v1 and v2 in the outputs of heads L14.H14 and L14.H18 at the position of y3 . (e) v3 in the layer-13 residual at the position of y3 .
Figure 28: The spaces at the position of y3 that contain the three-number directions, each in its top two principal components, coloured by the number its direction carries, with the smoothed mean position along that number (line) and the direction of the corresponding projection onto the plane (arrow). (a) The output of head L14.H14 with v1 . (b) The output of head L14.H18 with v2 . (c) The layer-13 residual with v3 .
Figure 29: Position recovery of a rank-one patch of each direction (rows) in each perturbation case (columns). Colours are clipped at 1 ; printed values are not.
Figure 30: Position recovery of each layer-14 head’s own output, interchanged at the runner-up’s position together with the full layer-13 residual there, for three value orders and the eight heads with the largest mean absolute effect. Dashed lines show the layer-13 patch alone for each order.
Figure 31: Position recovery of the patch of the span of v1,v2,v3 in the pre-MLP layer-14 residual at the position of y3 , by perturbed number, while parts of the network are frozen. (a) One sub-block’s output at the position of y3 frozen to its corrupted value. (b) The top k neurons by this task’s attribution ranking (solid) or k random neurons (dashed) frozen. (c) No neurons frozen, the 12 shared neurons of the two-number task frozen, and 12 random neurons frozen.
Figure 32: Position recovery of a rank-one direction in the layer-15 residual fitted in one case (rows) and evaluated in each case (columns), with the full-rank layer-15 patch in the last row, (a) at the position of y3 and (b) at the position of y2 . Colours are clipped at 2 ; printed values are not.
Figure 33: Top two principal components of the residual stream leaving layers 18 to 22 at the last token, on 1,500 clean three-number prompts, coloured by which operand holds the maximum.
Figure 34: (a) Position recovery of a full-rank patch of the residual leaving each of layers 18 to 22 at the last token, for three value orders. (b) Position recovery of each layer-20 head’s own output at the last token, for the eight heads with the largest effect and the same three orders. Dashed lines show a patch of the whole layer-20 attention sublayer for each order (Section B.4 ).
Figure 35: Accuracy on max(y1,…,yk) for 200 two-digit prompts per length, scored on the full decoded number (Section B.2 ): the intact model, the 12 shared neurons zeroed at every token, the same neurons zeroed only at the last token of the final operand (“last token” in the legend), and 12 random neurons zeroed at every token (two random sets, pooled).
Figure 36: IIA of the number-representation patches, scored on (a) the first generated token, the leading digit of r (Section B.2 ), and (b) the whole generated number: the answer decoded greedily for one token more than the operands’ digit count, counted correct when its first integer equals r . With y1 perturbed: the full layer-13 residual and u at the position of y1 , and v1 at the position of y2 , patched inside the output of L14.H14 as in the main text or in the pre-MLP layer-14 residual (hatched). With y2 perturbed: the full layer-13 residual and v2 at the position of y2 . The rank-one patches score the same on both, while the full-rank patches mostly generate a number other than r that starts with its leading digit.
Figure 37: IIA of the patches of the main-text shared-representation figure at the position of y2 , with y1 or y2 perturbed, scored on (a) the first generated token and (b) the whole generated number, as in Figure 36 : the full layer-13 residual, v2 , v1 , v1 and v2 as two separate rank-one patches, the rank-two (v1,v2) plane in the pre-MLP layer-14 residual, and the full residual leaving layer 14. Hatched bars are the two conditions that patch both directions. Only the full-rank patches with y2 perturbed lose IIA on the whole number.
Figure 38: IIA of the rank-one direction in the layer-15 residual at the position of y2 , fitted with y1 perturbed and evaluated in both cases, scored on the first generated token and on the whole generated number, as in Figure 36 .
y1 perturbed
y2 perturbed
patch
first digit
whole number
first digit
whole number
L13 full at y1
0.072
0.003
–
–
u
0.940
0.940
–
–
v1
0.460
0.460
0.000
0.000
v1 , pre-MLP
0.710
0.710
–
–
L13 full at y2
0.000
0.000
0.980
0.115
Appendix
Table 4: IIA on the first generated token and on the whole generated number for the patches of Figures 36 – 38 . Dashes mark cases in which a patch is not evaluated.
patch
perturbed
clean (y1,y2)
corrupted (y1,y2)
r
generated
L13 full at y1
y1
(93, 41)
(13, 41)
13
13
y1
(96, 46)
(20, 46)
20
26
y1
(69, 42)
(13, 42)
13
12
y1
(60, 38)
(15, 38)
15
12
y1
(99, 61)
(12, 61)
12
61
u
y1
(99, 61)
(12, 61)
12
12
Appendix
Table 5: Generated answers under the patches of Figure 36 , five per patch, chosen to include correct answers and mistakes: the perturbed number, the clean and corrupted operands, r , and the generated number. A patch succeeds when the model generates r .
patch
perturbed
clean (y1,y2)
corrupted (y1,y2)
r
generated
v1 & v2
y1
(99, 61)
(12, 61)
12
12
y2
(61, 99)
(61, 12)
12
12
y2
(44, 90)
(44, 11)
11
111
y2
(45, 79)
(45, 11)
11
111
y2
(55, 70)
(55, 33)
33
55
(v1,v2) plane
y1
(99, 61)
(12, 61)
12
12
Appendix
Table 6: Generated answers, as in Table 5 , under the remaining patches of Figures 37 and 38 .
College of Future Information Technology, Fudan University, Shanghai, 200433, China · Beijing National Research Center for Information Science and Technology, Tsinghua University, Beijing, 100084, China