Large language models can learn abstract rules from just a few in-context examples, but how their internal mechanisms activate as examples accumulate is not well understood. We trace a three-stage symbolic reasoning circuit (abstraction, induction, retrieval) across shot counts in three model families and find it is detectable and functional well before the model achieves high accuracy. Per-head causal contribution grows up to 8x from 1- to 10-shot, and cross-shot activation patching raises accuracy from 1% to 56% at 0-shot and 17% to 88% at 1-shot. Function vectors scaled and injected at 0-shot rescue accuracy up to 86%, largely substituting for the induction stage but depending critically on an intact downstream retrieval stage. The infrastructure for abstract rule-following is present in the weights before any demonstrations; in-context examples, function vectors, and related interventions appear to supply input to the same latent circuit.
Figures & tables
Figure 1 : The Symbolic Reasoning Circuit in Gemma 2-2B at 1 through 10-Shot. CMA Scores by layer and head by circuit stage across shot counts.
Figure 2 : Cross-Shot Patching Results (ABA). Accuracy of Gemma 2-2B when patching 10-shot donor activations into 1-shot (a) and 0-shot (b) target prompts. Full results for all models and both rules are in Appendix Figures 9 – 9 .
Figure 3 : Function-vector rescue at 0-shot, ABA.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Layers
Q heads
KV heads
dmodel
dhead
BOS
Gemma 2-2B ( Gemma Team, 2024 )
26
8
4
2304
256
yes
Llama 3.1-8B ( Grattafiori et al., 2024 )
32
32
8
4096
128
yes
Qwen 3-4B ( Yang et al., 2025a )
36
32
8
2560
128
no
Appendix
Table 1: Model architectures. All three use grouped-query attention (GQA). All experiments use bfloat16 precision and load the models via TransformerLens, which materializes one WK / WV per query head so attention-head indexing is uniform across architectures.
Stage
Patch position(s)
CMA condition
Symbolic Abstraction (SA)
all C positions: {b+6i+4:0≤i<n}
abstract
Symbolic Induction (SI)
final query position: b+6n+3
abstract
Retrieval (Ret)
final query position: b+6n+3
token
Appendix
Table 2: Patch positions per stage. SA writes to the C tokens of the demonstrations (so no SA evaluation is possible at n=0 ). SI and Retrieval write to the same query position but are distinguished by the contrastive condition.
Figure 4 : Cross-shot activation patching. Activations from 10-shot donors (middle) are patched into 1-shot (left) and 0-shot (right) targets (bottom), steering outputs to the correct rule-following completion. Yellow highlights indicate the SA head patch position (C tokens); blue indicates the SI and Retrieval head patch position (query token).
Condition
ABA
ABB
Baseline (no inject)
1.0
11.0
FV inject
86.0
88.5
FV inject + ablate SI
65.0
89.0
FV inject + ablate Ret
13.0
28.5
Random vector (norm-matched)
0.0
0.5
Appendix
Table 3: FV × per-stage ablation factorial at L=12 , s=25 . Accuracy (%), n=200 per cell.
Figure 5 : The Symbolic Reasoning Circuit in Llama 3.1-8B at 1 through 10-Shot. CMA Scores by layer by head, by circuit stage across shot counts.
Figure 6 : The Symbolic Reasoning Circuit in Qwen 3-4B at 1 through 10-Shot. CMA Scores by layer by head, by circuit stage across shot counts.
Stage
Shot
Top-5 (CMA score)
Gemma 2-2B
SA
1
L6H1 (0.50)
L8H1 (0.45)
L6H2 (0.38)
L3H1 (0.14)
L3H0 (0.14)
2
L6H1 (1.05)
L8H1 (0.88)
L6H2 (0.63)
L10H4 (0.47)
L3H1 (0.22)
4
L6H1 (1.73)
L8H1 (1.33)
L10H4 (1.01)
L6H2 (0.87)
L12H3 (0.58)
10
L6H1 (2.07)
L8H1 (2.03)
L10H4 (1.72)
L12H3 (1.30)
L6H2 (1.22)
SI
1
L14H0 (1.40)
L12H1 (0.73)
L15H0 (0.61)
L14H1 (0.31)
L13H0 (0.25)
Appendix
Table 4 : Top-5 attention heads per stage and shot count, ranked left-to-right by mean CMA score (in parentheses; averaged across pairs and across ABA/ABB rule directions). Bold marks persistent-core heads (top-5 at every shot count). Reading down a column shows reranking among the leading heads.
Stage
Top-5 core size
Top-10 core size
Top-5 core members
Gemma 2-2B
SA
3
6
L6H1, L6H2, L8H1
SI
4
6
L12H1, L14H0, L14H1, L15H0
Ret
3
7
L18H6, L20H7, L22H4
Llama 3.1-8B
SA
3
5
L2H27, L10H1, L14H19
Appendix
Table 5: Persistent cores: heads in the top- K at every shot count n∈{1,2,4,10} . Qwen’s top-5 is noisier at low shot counts but recovers at top-10.
1 ↔ 2
1 ↔ 4
1 ↔ 10
2 ↔ 4
2 ↔ 10
4 ↔ 10
Gemma 2-2B
SA
0.77
0.71
0.66
0.86
0.79
0.91
SI
0.81
0.77
0.72
0.86
0.81
0.91
Ret
0.84
0.79
0.78
0.94
0.86
0.87
Llama 3.1-8B
SA
0.85
0.58
0.54
0.67
0.62
0.81
Appendix
Table 6: Rank-Biased Overlap (RBO, p=0.9 , full-depth extrapolated) between per-head CMA rankings at every pair of shot counts. Each ranking averages across the 200 CMA pairs and across the ABA and ABB rule directions. Higher is more similar; for full rankings of unrelated lists RBO would tend to ≈1−p=0.1 .
Table 7: Cross-shot patching accuracy (%). Patching from a 10-shot donor into a 0/1-shot target, by stage. SA-only, SI-only, and Ret-only patch only the heads at that stage’s positions; All stages patches the union of all three. SI-shuffled uses a donor with the same rule but entirely disjoint tokens, patched at SI heads only. Random patches a matched-count random selection of non-circuit heads. n=200 prompts per cell. SA-only is identical to baseline at 0-shot because there are no in-context C positions to patch. Bold marks the All-stages column. Mean answer probability p tracks accuracy and is omitted for compactness.
Model
Shot
Succ.
Pred.
Gemma 2-2B
2
89.0%
17.6%
Gemma 2-2B
10
99.8%
76.4%
Llama 3.1-8B
2
98.4%
64.4%
Llama 3.1-8B
10
100%
100%
Appendix
Table 8: LSA accuracy (all three answer letters correct), 500 prompts per rule.
Figure 10 : LSA CMA heatmaps, Gemma 2-2B at 1, 2, and 10 shot, by stage. Same three-stage topology as ABA/ABB; L14H0 (SI) and L18H6 (Ret) dominate at all shot counts.