Human reasoning depends on how objects are related within propositions. \textit{How do relations organize the language representations of contextual contents?} We give an LLM a list of facts in its context (e.g., \emph{Alice eats an apple. Bob eats a pear.}) and measure how its hidden state changes when the question switches from what Alice eats to what Bob eats. Averaged over many lists, this change is a steering vector, which we call the \emph{ordinal vector}. It points to a fact by its \emph{order of mention}, the order in which the facts were stated in the context. We find that LLMs represent the fact a question asks about by its order of mention, not by the name the question contains. We state this as the \textit{ordinal addressing hypothesis}: each order of mention has a \emph{fact address} in the model's state, shared by all contexts, and a question moves the state to the fact address of the fact it asks about, while the context supplies what that fact says. Across Qwen, Gemma, and Llama, fact addresses are (1) \emph{ordered by mention}: query states are organized by the order of facts, not of names, even when one fact has multiple subjects; (2) \emph{steerable}: added to a question about the first fact of a new list, the ordinal vector makes the model answer with the second fact of that list; (3) \emph{low-rank}: they span a low-rank subspace in which the first-mentioned fact is the easiest to reach, surprisingly similar to human recall; and (4) \emph{emergent}: they are shared in late-middle layers, hold from 1.5B to 32B parameters, and form early in pretraining. Language models reach a stated fact by where it was mentioned, deepening our understanding of LLM reasoning.
Figures & tables
Figure 1: Language models address stated facts by order of mention. (a) Asked what Bob eats, the model’s state after the question points to the second fact, not to the name, and the answer is then read from that fact. (b) A real Qwen2.5-7B run: (1) the block-20 query change Δ between questions about fact 1 and fact 2, averaged over calibration contexts, is the ordinal vector v ; (2) added once in a new context, v moves the answer to fact 2’s animal.
Figure 2: The problem, the hypothesis, and its first test. (a) Definition 3.1 on the introduction’s example: each stated fact has an order of mention and a binding identity, and a question is answered from the binding identity at its order of mention. Swapping Bob and David changes the binding identities; reordering keeps the facts and changes pC ; the passive paraphrase changes neither. (b) Hypothesis 3.3 as a displacement field: every context shares the fact addresses a1,a2 , so the question moves each context by the same a2−a1 , or by −(a2−a1) when the two facts are reversed. (c) The same form measured in Qwen for eight test contexts, two arrangements at each of four offset levels (percentiles of the context-offset coordinate), on three axes fit on calibration word sets: along v , which people are named, and the context offset; units of ∥v∥ , with the offset axis stretched for display ( Appendix B ). (d) Cosines with v for new names (dark; mean of 64 name pairs) and between the ordinal vectors of reversed arrangements (purple), one dot per model, against the predictions of Proposition 3.4 ; small dots: eleven further models ( Appendix G ; open: base weights).
Figure 3: Query states are organized by order of mention, not by the queried person. (a–c) Query-dependent states of held-out word sets at each model’s main block, projected onto the top two principal directions of the four fact-address templates fit on eight other word sets; axis labels give the share of held-out query-dependent variance on each direction, and the unit is the fact-1 to fact-2 template distance. Each dot is one context under one question, colored by the queried fact’s order of mention; gray lines join a context’s fact-1 and fact-2 questions, the dark arrow is the predicted change a2−a1 , and contours enclose 50% and 85% of each order’s density. Panel headers: held-out R2 of fact-address and person templates.
Qwen
Gemma
Llama
A. Transfer to new contexts
Ordinal vector v
84.1 ± 2.0
53.9 ± 1.6
36.2 ± 1.6
Random
0.0
0.0
0.0
Ceiling
96.9 ± 1.3
96.4 ± 1.2
91.9 ± 1.9
B. Where the vector comes from
With facts ( v )
71.3 ± 4.0
41.7 ± 2.4
10.3 ± 2.0
Table 1: An ordinal vector from other contexts selects the designated fact in a new context. Designated answers (%), question unchanged; ± is half the 95% word-set bootstrap interval; shading is proportional to the rate. (A) 64 new word sets, 12,288 cases per model; Random: at most 0.01% . (B) 64 further word sets, 1,024 cases per row, all directions scaled to one length; the question and name rows are computed without the facts ( Appendix D ). Ceiling: the receiver’s own state for the designated question; it differs between panels because the word sets differ. (C) The receiver’s state replaced by a donor’s; 16 donors × 64 receivers; designated: the receiver’s fact at the order of mention of the other question.
Figure 4: Fact addresses are low-rank and ordered by mention. (a) Qwen’s held-out contexts in the address subspace fit on eight other word sets: query states (small dots, colored by the queried fact) and changes (gray); large dots: the four fact addresses, joined as a tetrahedron. (b) The eight fact addresses of eight-fact lists in their first two principal directions (unit: fact-1 to fact-2 distance); large marker: fact 1. (c) Distance between neighboring fact addresses, relative to facts 1–2, for four, six, and eight facts (light to dark). (d) Each address’s distance to its nearest neighbor (unit: fact-1 to fact-2 distance) against transfer to that fact ( Proposition 3.9 ), for 54 addresses; R2 after removing each model’s mean.
Figure 5: Through depth, the fact address is first shared across contexts, then held only in whole states, then replaced by the answer’s content. (a) The two interventions and three outcomes on one Gemma example, and each block classified over relative depth: shared address (the ordinal vector selects the receiver’s fact in at least 10% of cases), whole-state address (only donor states do, at least 10% ), and content (the donor’s first token reaches 50% ); blocks are classified one by one, so the first two can interleave; lower rows: other Qwen models ( Section G.1 ). (b–d) The three outcomes through depth, with the ceiling as reference, on the prompts of Table 1 C (one arrangement, two content assignments, four person changes; 16 donors × 64 receivers), where the ordinal vector selects 55.5% , 44.3% , and 25.8% at the main blocks, below Table 1 A over all arrangements; bands: 95% intervals resampling donor and receiver word sets. Shading: the depth window, blocks where the ordinal vector selects at least 10% .
Figure 6: In pretraining, questions switch from the named person to order of mention before the model answers reliably, and the full depth profile forms late. Ten checkpoints of OLMo-2-7B, from 5B training tokens to the released model ( final , after mid-training). (a) Stages as in Figure 5 a, with each threshold applied to the gain over the unedited answer; all three stages first co-occur at 839B. (b) Fact-address R2 (peak over blocks) and the cosines of Proposition 3.4 at block 16. (c) Designated answers at block 16; dashed: unedited accuracy; dotted: random directions.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Inside one word set, the query state passes from the named person to the fact address to the requested content. Held-out R2 of query-dependent states under templates for the queried person, the queried fact’s order of mention, and the requested content, each fit on the other 15 contexts of the same word set and pooled over 16 calibration word sets. Shading: first to last block where the ordinal vector selects at least 10% of designated answers ( Figure 5 ).
Figure 8: Queries over two-place facts address a fact and an argument slot. (a) Designated answers for changes of fact, of argument slot, and of both, against the ceiling (gray ticks) and random directions ( × ), on 64 new active contexts. (b) Held-out R2 under fact-by-slot, additive (fact + slot), fact-only, slot-only, and person templates. (c) Cosines of fact and of slot changes across relations. In (b) and (c), small markers: individual relations; large markers: their mean.
Model
Change
Original pair
New order of mention
Recalibrated
Ceiling
Random
No facts
Qwen
Person
9.0 ± 1.1
69.7 ± 2.3
80.9 ± 2.5
90.2 ± 1.7
0.0
0.0
Relation
22.6 ± 1.6
61.6 ± 2.1
68.5 ± 1.9
77.3 ± 1.9
5.7
7.0
Gemma
Person
17.6 ± 0.9
40.8 ± 1.4
61.1 ± 2.1
79.2 ± 1.7
0.1
0.1
Relation
5.5 ± 1.3
19.9 ± 2.9
41.5 ± 2.8
73.6 ± 1.6
0.8
0.7
Llama
Person
6.2 ± 1.4
15.7 ± 2.8
16.4 ± 2.8
29.8 ± 3.7
0.0
0.0
Relation
0.7 ± 0.5
0.6 ± 0.3
2.8 ± 1.1
11.8 ± 2.1
0.3
0.3
Appendix
Table 2: When facts move, the ordinal vector must follow their orders of mention. Designated answers (%) under two held-out arrangements, 64 test word sets and 2,048 cases per entry; ± is half the 95% word-set bootstrap interval; bold: the new-order-of-mention edit; ceiling: the receiver’s own state for the designated question; no facts: the direction computed from the question alone.
Figure 9: The same ordinal vector acts under every relation. (a) Designated answers by receiving relation: an ordinal vector calibrated on the same relation (filled), on each of the five other relations at its own length (open), and random directions ( × ); 64 test word sets, 3,072 cases per point. Not plotted: the eats vector rescaled to the receiving relation’s change length comes within five points of the filled marker for every relation (paired differences −4.0 to +4.6 points). (b) Held-out R2 of query states under fact-address (filled) and person (open) templates, one point per relation. (c) Cosine between ordinal vectors fit separately on two relations, one point per relation pair, averaged over pairs of orders of mention; gray bars are the cosine between mismatched ordinal vectors.
Figure 10: The ordinal vector follows order of mention, not token distance. (a) Designated answers by length condition for ordinal vectors calibrated on uniform (filled) or mixed-length (open) contexts, with the ceiling (gray ticks) and random directions ( × ). (b) Answers naming the fact at the calibrated order minus those naming the fact at the calibrated absolute or relative distance, on edits where the two facts differ from each other and from the unedited answer; markers as in (a), with 95% intervals over 64 word sets.
Figure 11: Transfer falls with the designated fact’s order of mention and is highest for the first fact. Designated answers for the ordinal vector by the designated fact’s order of mention, in lists of four, six, and eight facts, with word-set bootstrap 95% intervals (bands, mostly narrower than the markers); dashed lines are the ceiling. Random directions give at most 0.1% at every length.
Edit
Qwen
Gemma
Llama
Joint change vJ
68.5
57.1
26.0
Same-start sum vE+vR
43.8
44.0
18.8
Sum rescaled to ∥vJ∥
35.3
31.7
14.1
Ceiling (own state for the joint question)
87.9
90.7
56.7
Joint minus sum
24.7 ± 1.3
13.1 ± 1.4
7.2 ± 0.9
Joint minus rescaled sum
33.2 ± 1.5
25.4 ± 1.4
11.9 ± 1.0
Appendix
Table 3: The joint change selects the joint target more often than the sum of its parts. Designated answers (%), 64 test word sets, 4,096 cases per edit; ± : half the 95% paired word-set bootstrap interval, in points; bold: the joint change.
Figure 12: Fact addresses predict the person–relation interaction. Distribution, over 3,328 held-out contexts per model, of the cosine between the measured interaction I of the query-dependent states ( Equation 4 ) and its prediction from fact-address templates and the context’s orders of mention; means 0.47 , 0.38 , and 0.46 in Qwen, Gemma, and Llama. Dashed line: the mean cosine when the order labels are shuffled ( −0.02 ).
Figure 13: The fact-address geometry follows the stages of Figure 5 . Query-state geometry at every block: R2 of fact-address and person templates (leave-one-word-set-out over the 16 calibration word sets, 256 contexts), and each query state’s alignment with its answer’s first-token output direction (128 test contexts). Shading: the depth window of Figure 5 .
At block 16
Scored by stated content (%)
Cosine
Tokens
Rule block
Unedited
v
Random
Ceiling
New names
Reversed
Peak R2 (block)
5B
22
1.3
1.2
1.2
1.2
0.09
0.98
0.00 (31)
17B
22
24.9
16.3
16.0
16.7
0.37
0.25
0.29 (11)
42B
14
50.0
25.1
14.6
31.5
0.70
− 0.61
0.52 (13)
105B
22
63.9
23.7
11.3
27.6
0.76
− 0.74
0.61 (13)
Appendix
Table 4: OLMo-2-7B checkpoints. Rule-chosen block; at block 16: unedited accuracy and designated answers (%) for the ordinal vector (bold), a random direction and the ceiling, scored by the single stated content an answer names, 64 test word sets and 12,288 cases per edit; new-name and reversed-order cosines of Proposition 3.4 ; and the peak leave-one-word-set-out fact-address R2 over blocks (its block). final : the released model after stage-2 mid-training.
Designated answers (%)
Change cosine
R2
Model
Main block
Unedited
v
Random
Ceiling
New names
Reversed
Fact address
Person
Window
Qwen2.5-1.5B
18/28
98.6
72.0
0.1
95.7
0.97
− 0.96
0.76
− 0.02
18–22
Qwen2.5-3B
26/36
96.6
48.0
0.0
66.1
0.92
− 0.85
0.65
− 0.02
26–30
Qwen2.5-7B †
20/28
99.9
84.1
0.0
96.9
0.93
− 0.93
0.66
− 0.01
18–22
Qwen2.5-14B
34/48
99.9
44.4
0.0
65.2
0.92
− 0.92
0.46
− 0.01
32–34
Qwen2.5-32B
48/64
100.0
55.1
0.0
74.7
0.94
− 0.94
0.51
− 0.01
46–50
Appendix
Table 5: Ordinal addressing holds from 1.5B to 32B parameters. Every model runs the transfer test of Table 1 A (64 new word sets, 12,288 cases), the cosines of Figure 2 d, the fact-address and person templates of Section C.1 , and the depth curves of Figure 5 at every second block. Main block: block index / number of blocks, chosen on block-selection word sets by the rule of Appendix A . Unedited: correct answers without an edit. Window: first to last block at which the ordinal vector selects at least 10% , as in Figure 5 (8 donors × 32 receivers). † The main-text model; its block was chosen before this sweep ( Appendix A ), and its window comes from the every-block curves of Figure 5 .
Designated answers (%)
Change cosine
Model
Weights, prompt
Block
Unedited
v
Random
Ceiling
New names
Reversed
Fact- address R2
Qwen2.5-1.5B
base, plain
18
87.7
67.3
1.7
82.0
0.97
− 0.95
0.78
instruct, plain
20
98.7
69.6
0.1
95.3
0.95
− 0.94
0.63
instruct, chat
18
98.6
72.0
0.1
95.7
0.97
− 0.96
0.76
Qwen2.5-3B
base, plain
26
95.5
80.0
0.1
94.1
0.94
− 0.87
0.69
instruct, plain
26
95.7
72.2
0.0
90.3
0.93
− 0.88
0.65
Appendix
Table 6: Base models address facts the same way. Plain: the instruction, facts, and question followed by Answer: , with no chat template, identical for base and instruct weights; chat: the instruct model’s native template. Each row chooses its own main block by the rule of Appendix A and runs the tests of Table 5 . ∗ Gemma-3-4B base answers in sentences, so its answers are rescored by the single stated content they name. ‡ Main-text Gemma run ( Table 1 A, Figure 2 d, Section C.1 ).
Ordinal vector calibrated on
Model
Receiver format
Unedited
Lists (main)
Lists (new word sets)
Prose
Random
Ceiling
Qwen
prose
99.7
44.5 ± 1.8
44.0 ± 1.8
40.2 ± 1.5
0.0
92.8 ± 1.6
list
99.9
80.5 ± 2.1
78.1 ± 2.2
41.1 ± 1.7
0.0
95.4 ± 1.5
Gemma
prose
99.2
42.4 ± 1.2
43.4 ± 1.2
32.3 ± 1.2
0.0
96.3 ± 0.9
list
99.7
52.2 ± 1.4
52.8 ± 1.4
24.1 ± 1.2
0.0
96.5 ± 0.9
Llama
prose
99.6
14.0 ± 1.1
10.2 ± 0.9
8.5 ± 0.9
0.0
83.1 ± 1.5
Appendix
Table 7: Ordinal addressing carries over to prose. Top: designated answers (%) on 64 new word sets, 12,288 cases per entry, each edit added once at the main block. Receivers state the four facts either as a paragraph ( prose ) or as the list of the main text, with the same people, contents, arrangements, and questions. The ordinal vectors are calibrated on the main-text lists (the v of Table 1 A), on lists of 16 new word sets, or on paragraphs of the same 16 word sets. Bottom: at the main block, leave-one-word-set-out R2 of fact-address and person templates on paragraphs and on the matched lists (the list R2 uses this experiment’s word sets, not those of Section C.1 ); mean cosine between matching prose and list fact-address templates and between matching prose and list ordinal vectors; fraction of the squared norm of query-dependent prose states inside the span of the main-text ordinal vectors, against random subspaces of the same dimension.
University of Science and Technology of China Hefei, China · Singapore Management University Singapore, Singapore · The University of Tokyo Tokyo, Japan +1