When Models Don't Manipulate Manifolds: The Geometry of a Comparison Task
Organizations: School of Engineering and Applied Science, Harvard University · Kempner Institute, Harvard University · Goodfire AI
Abstract
One of the current premises of mechanistic interpretability research is that detailed accounts of the geometry of neural network representations can tell us how models perform computations, and how to effectively intervene on them. While low dimensional manifolds have been observed for multiple concepts in the literature (e.g. numbers encoded on helices, days of the week on a circle, ...), with structure believed to reflect properties of data and tasks, the extent to which models rely on them for computation, and how they manipulate them, remains unclear. We characterize precisely the geometry of computation in a number-comparison task, as an abstraction of comparison for decision making, and how models utilize geometry in an elegant fashion to implement it. Specifically, we study the causal geometry of number comparison in Qwen2.5-7B-Instruct, a capable and widely studied open-weight model, and find Qwen largely uses linear representations of numbers despite the presence of curved geometry. To compare two numbers, the model first encodes each number along a vector and adds the two representations using attention and the residual connection, bringing them into a shared space in the residual stream. Then, the model uses MLP neurons to compare the pair of numbers on local regions in this shared space, which correspond to smaller intervals of input numbers, and combines these to obtain the position of the maximum. In fact, this reliance on linear representations for comparison also persists when the model compares three numbers. Our findings demonstrate that the manifold hypothesis can co-exist with linear representations: while concepts that are ordered may have manifold structure in representations, the model may use an underlying linear structure of the concept in certain computations.
Figures & tables
Appendix figures & tables39 assets
Supplementary material from the paper’s appendix.
Appendix
| quantity | value | |
| populations | evaluation / fitting examples | / |
| clouds (L13 two-number / L14 / three-number) | / / | |
| operand counts in the accuracy sweeps, prompts per | ; | |
| base random seed | ||
| learned directions | Adam steps, learning rate | , |
| minibatch (gradients accumulated) | – |
| direction | position | site | obtained by | |
|---|---|---|---|---|
| L13 residual | DAS, perturbed | |||
| output of L14.H14 | DAS, perturbed | |||
| L13 residual | PC | |||
| comparison flag | L15 residual | DAS, each case | ||
| , | , | L13 residual | DAS, cases , | |
| , | outputs of L14.H14, L14.H18 | DAS, cases , |
| intervention | clean | max | corrupted | max | patched answer |
|---|---|---|---|---|---|
| at | (99, 61) | 99 | (12, 61) | 61 | 12 |
| at | (70, 55) | 70 | (33, 55) | 55 | 33 |
| full at | (61, 99) | 99 | (61, 12) | 61 | 12 |
| full at | (56, 97) | 97 | (56, 27) | 56 | 27 |
| full at | (55, 70) | 70 | (55, 33) | 55 | 30 |
| full at | (71, 89) | 89 | (71, 52) | 71 | 59 |
| perturbed | perturbed | |||
|---|---|---|---|---|
| patch | first digit | whole number | first digit | whole number |
| L13 full at | 0.072 | 0.003 | – | – |
| 0.940 | 0.940 | – | – | |
| 0.460 | 0.460 | 0.000 | 0.000 | |
| , pre-MLP | 0.710 | 0.710 | – | – |
| L13 full at | 0.000 | 0.000 | 0.980 | 0.115 |
| patch | perturbed | clean | corrupted | generated | |
|---|---|---|---|---|---|
| L13 full at | (93, 41) | (13, 41) | 13 | 13 | |
| (96, 46) | (20, 46) | 20 | 26 | ||
| (69, 42) | (13, 42) | 13 | 12 | ||
| (60, 38) | (15, 38) | 15 | 12 | ||
| (99, 61) | (12, 61) | 12 | 61 | ||
| (99, 61) | (12, 61) | 12 | 12 |
| patch | perturbed | clean | corrupted | generated | |
|---|---|---|---|---|---|
| & | (99, 61) | (12, 61) | 12 | 12 | |
| (61, 99) | (61, 12) | 12 | 12 | ||
| (44, 90) | (44, 11) | 11 | 111 | ||
| (45, 79) | (45, 11) | 11 | 111 | ||
| (55, 70) | (55, 33) | 33 | 55 | ||
| plane | (99, 61) | (12, 61) | 12 | 12 |