Understanding large language model (LLM) reasoning requires looking beyond behavioral performance to examine how reasoning ability is reflected in internal organization. Inspired by neuroscience findings linking higher intelligence to stronger small-world organization in functional brain networks, we investigate small-world connectivity as a structural signature of LLM reasoning. We construct functional graphs from attention-head activation similarities and find that a higher small-world index (SWI), capturing local clustering and short global paths, consistently correlates with better fluid reasoning performance across models and training checkpoints. Since local clustering is central to small-world organization, we further examine how heads important for model performance connect within and across communities. We find that these heads tend to have a larger share of connection weight within their own communities (high core scores) and a more concentrated weight distribution across communities (low bridge scores). These observations motivate the hypothesis that high core and low bridge scores serve as structural indicators of head importance for reasoning capability. We validate this hypothesis through pruning, introducing Small-World Allocation (SWA), a hierarchical sparsity allocation method guided by these scores. Across six LLMs, SWA better preserves small-world organization and model performance than competing allocation strategies, reducing WikiText perplexity by up to 20%. Together, these findings identify small-world functional connectivity as a measurable signature of LLM reasoning performance, offering a structural perspective that complements behavioral evaluation.
Figures & tables
Figure 1: We extract activations from attention heads across Transformer layers and compute pairwise similarity to construct a head-level functional graph. The resulting graph exhibits small-world structure, characterized by high clustering and efficient short-path communication. We show that this structural property correlates with model capability on fluid intelligence benchmarks.
Figure 2
Head property
Top 25%
Remaining 75%
p -value
Low Bridge Score
89.3%
36.9%
3.61×10−34
High Core Score
49.4%
23.8%
5.95×10−9
Table 1: Prevalence of low bridge scores ( b ) and high core scores ( c ) in Llama-3.2-3B on DRE-Bench, comparing the top 25% of heads ranked by gradient-based importance with the remaining 75%. Low bridge and high core scores are defined as b≤median(b) and c>median(c) , respectively, with medians computed across all heads. Two-sided p -values test the null hypothesis of equal prevalence between the two groups and are adjusted for multiple comparisons using Holm’s method ( Holm, 1979 ) . Smaller p -values indicate stronger evidence that the two groups differ in the proportion of heads exhibiting the corresponding structural feature.
Model
SparseGPT
+ FARMS
+ ATP
+ SWA (Head)
+ SWA (Layer)
+ SWA
Qwen3 14B
0.8000 ±0.76
0.8031 ±0.09
0.8797 ±0.37
0.9867 ±0.23
1.057 ±0.14
1.1386 ±0.40
Qwen3 8B
0.6509 ±0.64
0.7982 ±0.90
0.8432 ±0.30
0.8639 ±0.22
0.9053 ±0.15
0.9196 ±0.44
Qwen3 4B
0.5400 ±0.10
0.5516 ±0.81
0.6555 ±0.14
0.6449 ±0.06
0.7107 ±0.25
0.7212 ±0.58
Llama3 8B
0.8085 ±0.05
0.8229 ±0.28
0.8300 ±0.68
0.8634 ±0.64
0.8943 ±0.22
0.9040 ±0.45
Llama3.2 3B
0.8372 ±0.21
0.8380 ±0.27
0.8530 ±0.14
0.8468 ±0.26
0.8777 ±0.09
0.8867 ±0.24
Llama3.2 1B
0.5138 ±0.50
0.5408 ±0.40
0.5567 ±0.54
0.5943 ±0.93
0.6641 ±0.67
0.6740 ±0.43
Table 2: We report the small-world index (SWI ↑ ) of head-head functional graphs after applying SparseGPT and its variants at sparsity 0.5. SWA (Head) and SWA (Layer) denote head-only and layer-only allocation strategies, respectively. Higher SWI indicates better preservation of small-world functional organization under pruning. Results are averaged over three seeds and reported as mean ±std . The best result for each model is bolded .
Model
Size
Sparsity
Wanda
SparseGPT
Base
+ FARMS
+ ATP
+ SWA
Base
+ FARMS
+ ATP
+ SWA
Llama 3.2
1B
0.5
36.25 ±0.18
33.90 ±0.02
24.95 ±0.04
21.92 ±0.05
21.54 ±0.29
21.14 ±0.56
20.56 ±0.34
20.47 ±0.15
0.7
415.60 ±32.1
380.20 ±20.1
264.10 ±21.7
256.67 ±18.00
157.20 ±3.2
136.80 ±5.3
106.60 ±3.2
105.00 ±3.6
3B
0.5
19.71 ±0.02
16.45 ±0.09
14.42 ±0.02
13.32 ±0.12
23.98 ±2.15
15.92 ±1.85
13.53 ±0.08
11.27 ±0.98
0.7
177.80 ±5.6
132.6 ±5.3
125.1 ±0.8
118.67 ±2.32
70.70 ±5.07
67.67 ±2.93
51.51 ±3.83
49.30 ±0.93
Llama 3
8B
0.5
11.86 ±0.02
11.12 ±0.03
11.50 ±0.05
10.14 ±0.03
11.69 ±0.13
11.02 ±0.87
10.72 ±0.15
10.15 ±0.08
Table 3: Comparison of WikiText Perplexity ( ↓ ) . We evaluate Base pruning methods (Wanda and SparseGPT) and their variants augmented with FARMS, ATP, and SWA at sparsity 0.5/0.7. Results are averaged over three seeds and reported as mean ±std . The best result in each setting is bolded . SWA reduces perplexity and achieves the best values across all the models and sparsity levels.
Variant
SparseGPT
Wanda
0.5
0.7
0.5
0.7
SWA
12.10
20.10
11.90
52.56
SWA(Layer)
+3.00
+10.21
+0.08
+5.73
SWA(Head)
+3.69
+19.97
+0.17
+44.99
Table 4: Ablation study on hierarchical structural deletion. We report the absolute WikiText perplexity (PPL ↓ ) of SWA and the increase in PPL for each ablated variant relative to SWA on Qwen3 8B at sparsity levels of 0.5 and 0.7. SWA(Layer) uses only layer-level scores, and SWA(Head) uses only head-level scores.
Model
Base
GSM8K
ARC-C
MMLU
Llama3 8B
55.94 ±3.23
42.60 ±1.22
41.10 ±1.35
40.75 ±1.07
Qwen3 4B
50.93 ±1.62
40.42 ±1.12
38.90 ±1.23
39.89 ±2.35
Llama3.2 3B
70.70 ±5.07
50.54 ±1.30
49.69 ±2.69
48.95 ±1.20
Llama3.2 1B
157.20 ±3.20
109.2 ±6.20
110.9 ±2.60
105.28 ±6.62
Table 5: Effect of the graph-construction benchmark on SWA. We report WikiText perplexity ( ↓ ) for SparseGPT at 70% sparsity, comparing the base method with SWA using functional graphs built from GSM8K, ARC-C, or MMLU.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Training step
Tokens seen
Level 1
Level 2
NLL/token ↓
SWI ↑
NLL/token ↓
SWI ↑
4,000
8B
0.5089 ±0.16
0.6760 ±0.85
0.8255 ±0.25
0.4587 ±0.67
11,000
23B
0.4677 ±0.20
0.8008 ±0.13
0.7020 ±0.68
0.5574 ±0.80
30,000
63B
0.4344 ±0.28
0.8035 ±0.12
0.7007 ±0.52
0.5599 ±0.54
65,000
136B
0.2250 ±0.30
0.8135 ±0.19
0.6411 ±0.34
0.8100 ±0.32
71,000
149B
0.1717 ±0.07
0.8320 ±0.47
0.4026 ±0.18
0.8469 ±0.33
Appendix
Table 6: Relationship between DRE-Bench performance and small-world organization across Pythia-6.9B training checkpoints. We report the mean per-token negative log-likelihood over the answer tokens (NLL/token; lower is better) and SWI of the head-head functional graph on DRE-Bench Level 1 and Level 2 tasks. Results are averaged over three seeds and reported as mean ±std . Across training checkpoints, lower NLL is associated with higher SWI on both levels.
Model
Level 1
Level 2
NLL/token ↓
SWI ↑
NLL/token ↓
SWI ↑
Qwen3 14B
0.0822 ±0.01
1.8160 ±0.27
0.2410 ±0.03
1.7020 ±0.12
Qwen3 8B
0.0949 ±0.01
1.5850 ±0.28
0.2814 ±0.02
1.6000 ±0.07
Llama3 8B
0.0980 ±0.01
1.3750 ±0.19
0.2993 ±0.02
1.5935 ±0.11
Qwen3 4B
0.1142 ±0.02
1.2180 ±0.26
0.3397 ±0.05
1.2722 ±0.18
Llama3.2 3B
0.1377 ±0.07
1.1933 ±0.33
0.3412 ±0.10
1.2100 ±0.17
Appendix
Table 7: Relationship between DRE-Bench performance and small-world organization. We report the mean per-token negative log-likelihood over the answer tokens (NLL/token; lower is better) and SWI of the head-head functional graph on DRE-Bench Level 1 and Level 2 tasks. Results are averaged over three seeds and reported as mean ±std . Across models, lower NLL is associated with higher SWI.
Model
Size
Sparsity
Wanda
SparseGPT
Base
+ FARMS
+ ATP
+ SWA
Base
+ FARMS
+ ATP
+ SWA
Llama 3.2
1B
0.5
34.91 ±3.67
35.21 ±2.28
37.16 ±0.82
38.89 ±1.03
39.57 ±2.67
40.33 ±2.41
41.26 ±0.23
44.67 ±1.75
0.7
25.86 ±3.62
27.29 ±4.75
29.86 ±4.65
32.41 ±4.81
30.85 ±4.67
29.38 ±1.33
35.29 ±0.22
36.46 ±0.34
3B
0.5
40.53 ±1.41
42.71 ±1.22
44.57 ±1.39
45.80 ±1.87
45.50 ±0.36
46.67 ±0.50
46.05 ±1.41
49.91 ±0.31
0.7
40.16 ±2.41
41.50 ±1.63
40.91 ±4.74
43.04 ±0.87
40.42 ±1.16
42.25 ±0.85
41.69 ±4.46
43.57 ±1.11
Llama 3
8B
0.5
52.12 ±1.25
53.49 ±2.17
52.43 ±1.32
55.13 ±1.52
48.53 ±2.24
50.28 ±1.30
53.48 ±0.61
55.61 ±0.85
Appendix
Table 8: Comparison of Zero-shot Accuracy ( ↑ ). We evaluate Base pruning methods (Wanda and SparseGPT) and their variants augmented with FARMS, ATP, and SWA at sparsity 0.5/0.7. Results are averaged over three seeds and reported as mean ±std . The best result in each setting is bolded . SWA improves zero-shot accuracy and achieves the strongest performance over all the models and sparsity levels.
Method
Move
Change Color
Copy
Mirror
Fill Internal
Scale
Qwen3 8B (Non-Pruned)
36.00 ±4.11
23.14 ±4.61
12.18 ±4.12
84.66 ±3.20
6.07 ±1.99
83.33 ±5.87
SparseGPT
44.21 ±9.20
48.87 ±10.40
60.99 ±10.67
94.76 ±8.21
34.25 ±4.67
88.91 ±8.03
SparseGPT + FARMS
46.22 ±8.21
50.77 ±15.85
56.60 ±12.80
96.66 ±5.48
33.11 ±5.11
88.06 ±8.95
SparseGPT + ATP
43.09 ±9.01
47.73 ±8.66
56.79 ±12.48
93.56 ±9.63
32.66 ±5.81
89.49 ±7.47
SparseGPT + SWA
40.29 ±9.25
46.28 ±9.62
56.11 ±11.62
90.21 ±8.89
28.69 ±4.28
85.10 ±7.28
Appendix
Table 9: ARAOC task mismatch rate at sparsity 0.5. We report mismatch rate (Not M ↓ ) across six tasks. Results are averaged over three runs and reported as mean ±std . The best result in each setting is bolded .
Large language models (LLMs) excel at multi-step reasoning but incur substantial inference cost. We introduce Causal Attribution Pruning (CAP), a training-free method that identifies critical attention heads by measuring their causal impact on reasoning tasks and uses these head-level scores to guide fine-grained weight pruning. For each attention head, CAP estimates the expected performance degradation when the head is masked during forward passes on a small calibration set of reasoning problems. These causal scores are then converted into weight-level importance values for the corresponding projection matrices. Unlike magnitude-only or activation-based criteria, CAP's interventional measurement directly captures each head's functional contribution, yielding relative accuracy gains of up to 61% over Wanda on ARC-Challenge at 20% sparsity. We evaluate CAP on GSM8K, StrategyQA, and ARC-Challenge using Llama-3-8B-Instruct and Mistral-7B-Instruct at 10%, 20%, and 50% sparsity. At moderate sparsity (10-20%), CAP improves over Wanda in most model-benchmark configurations. with especially large gains on ARC-Challenge for Llama-3. Our results suggest that attention-head-level causal attribution can better preserve reasoning performance on downstream benchmarks than correlational pruning criteria at equivalent sparsity, while remaining limited by coarse MLP attribution at 50% sparsity.
Amogh Sheth, Biruk Assefa, Yi Wen Huang +2
Edison Academy Magnet School · State University of New York College at Plattsburgh · The University of Texas at Austin +2
Recent studies have shown that Large Language Models (LLMs) can achieve strong reasoning performance by incorporating functional symbolic representations that abstractly describe graph traversal algorithms and step-by-step reasoning in few-shot learning settings. However, it remains unclear how LLMs genuinely understand the abstract meaning of each reasoning step and the overall algorithm from only a limited number of demonstrations. This work aims to localize the attention heads responsible for individual reasoning steps and characterize the types of information transferred among them. We first align constituent reasoning steps with their corresponding token logits under a symbolic-aided Chain-of-Thought (CoT) prompting framework. Our analysis shows that token positions that steer the reasoning process are associated with low confidence scores caused by constraints on satisfying reasoning behavior patterns in demonstrations. We then adopt causal mediation analysis techniques to identify the attention heads responsible for these patterns. In addition, our findings indicate that LLMs retrieve factual and rule-based information for individual sub-reasoning tasks through specialized attention heads (approximately 3% total heads), whereas higher layers predominantly facilitate information integration and the emergence of global reasoning strategies (e.g., graph traversal algorithms) that coordinate multiple intermediate reasoning steps to solve the overall task.
Phuong Minh Nguyen, Tien Huu Dang, Naoya Inoue
Japan Advanced Institute of Science and Technology
Large language models (LLMs) increasingly exhibit behaviors suggesting awareness of their evaluation context, often adapting their reasoning strategies in benchmark settings. Prior work has shown that such evaluation awareness can distort performance measurements; however, it remains unclear whether this phenomenon reflects a single behavioral artifact or a deeper internal structure within the model. We propose that LLMs maintain a decomposable space of functional metacognitive states: internal variables encoding factors such as evaluation awareness, self-assessed capability, perceived risk, computational effort allocation, audience expertise adaptation, and intentionality. Through residual stream analysis across multiple reasoning models, we demonstrate that these states are linearly decodable from internal activations and exhibit distinct layer-wise profiles. Moreover, by steering model activations along probe-derived directions, we show that each functional metacognitive state causally modulates reasoning behavior in dissociable ways, affecting verbosity, accuracy, and safety-related responses across tasks. Our findings suggest that benchmark performance reflects not only task competence but also the activation of specific functional metacognitive states. We argue that understandi ng and controlling these internal states is essential for reliable evaluation and deployment of reasoning models, and we provide a mechanistic framework for studying functional m etacognition in artificial systems. Our code and data are publicly available at https://github.com/xlands/meta-cognition.