Understanding large language model (LLM) reasoning requires looking beyond behavioral performance to examine how reasoning ability is reflected in internal organization. Inspired by neuroscience findings linking higher intelligence to stronger small-world organization in functional brain networks, we investigate small-world connectivity as a structural signature of LLM reasoning. We construct functional graphs from attention-head activation similarities and find that a higher small-world index (SWI), capturing local clustering and short global paths, consistently correlates with better fluid reasoning performance across models and training checkpoints. Since local clustering is central to small-world organization, we further examine how heads important for model performance connect within and across communities. We find that these heads tend to have a larger share of connection weight within their own communities (high core scores) and a more concentrated weight distribution across communities (low bridge scores). These observations motivate the hypothesis that high core and low bridge scores serve as structural indicators of head importance for reasoning capability. We validate this hypothesis through pruning, introducing Small-World Allocation (SWA), a hierarchical sparsity allocation method guided by these scores. Across six LLMs, SWA better preserves small-world organization and model performance than competing allocation strategies, reducing WikiText perplexity by up to 20%. Together, these findings identify small-world functional connectivity as a measurable signature of LLM reasoning performance, offering a structural perspective that complements behavioral evaluation.
Figures & tables
Figure 1: We extract activations from attention heads across Transformer layers and compute pairwise similarity to construct a head-level functional graph. The resulting graph exhibits small-world structure, characterized by high clustering and efficient short-path communication. We show that this structural property correlates with model capability on fluid intelligence benchmarks.
Figure 2
Head property
Top 25%
Remaining 75%
p -value
Low Bridge Score
89.3%
36.9%
3.61×10−34
High Core Score
49.4%
23.8%
5.95×10−9
Table 1: Prevalence of low bridge scores ( b ) and high core scores ( c ) in Llama-3.2-3B on DRE-Bench, comparing the top 25% of heads ranked by gradient-based importance with the remaining 75%. Low bridge and high core scores are defined as b≤median(b) and c>median(c) , respectively, with medians computed across all heads. Two-sided p -values test the null hypothesis of equal prevalence between the two groups and are adjusted for multiple comparisons using Holm’s method ( Holm, 1979 ) . Smaller p -values indicate stronger evidence that the two groups differ in the proportion of heads exhibiting the corresponding structural feature.
Model
SparseGPT
+ FARMS
+ ATP
+ SWA (Head)
+ SWA (Layer)
+ SWA
Qwen3 14B
0.8000 ±0.76
0.8031 ±0.09
0.8797 ±0.37
0.9867 ±0.23
1.057 ±0.14
1.1386 ±0.40
Qwen3 8B
0.6509 ±0.64
0.7982 ±0.90
0.8432 ±0.30
0.8639 ±0.22
0.9053 ±0.15
0.9196 ±0.44
Qwen3 4B
0.5400 ±0.10
0.5516 ±0.81
0.6555 ±0.14
0.6449 ±0.06
0.7107 ±0.25
0.7212 ±0.58
Llama3 8B
0.8085 ±0.05
0.8229 ±0.28
0.8300 ±0.68
0.8634 ±0.64
0.8943 ±0.22
0.9040 ±0.45
Llama3.2 3B
0.8372 ±0.21
0.8380 ±0.27
0.8530 ±0.14
0.8468 ±0.26
0.8777 ±0.09
0.8867 ±0.24
Llama3.2 1B
0.5138 ±0.50
0.5408 ±0.40
0.5567 ±0.54
0.5943 ±0.93
0.6641 ±0.67
0.6740 ±0.43
Table 2: We report the small-world index (SWI ↑ ) of head-head functional graphs after applying SparseGPT and its variants at sparsity 0.5. SWA (Head) and SWA (Layer) denote head-only and layer-only allocation strategies, respectively. Higher SWI indicates better preservation of small-world functional organization under pruning. Results are averaged over three seeds and reported as mean ±std . The best result for each model is bolded .
Model
Size
Sparsity
Wanda
SparseGPT
Base
+ FARMS
+ ATP
+ SWA
Base
+ FARMS
+ ATP
+ SWA
Llama 3.2
1B
0.5
36.25 ±0.18
33.90 ±0.02
24.95 ±0.04
21.92 ±0.05
21.54 ±0.29
21.14 ±0.56
20.56 ±0.34
20.47 ±0.15
0.7
415.60 ±32.1
380.20 ±20.1
264.10 ±21.7
256.67 ±18.00
157.20 ±3.2
136.80 ±5.3
106.60 ±3.2
105.00 ±3.6
3B
0.5
19.71 ±0.02
16.45 ±0.09
14.42 ±0.02
13.32 ±0.12
23.98 ±2.15
15.92 ±1.85
13.53 ±0.08
11.27 ±0.98
0.7
177.80 ±5.6
132.6 ±5.3
125.1 ±0.8
118.67 ±2.32
70.70 ±5.07
67.67 ±2.93
51.51 ±3.83
49.30 ±0.93
Llama 3
8B
0.5
11.86 ±0.02
11.12 ±0.03
11.50 ±0.05
10.14 ±0.03
11.69 ±0.13
11.02 ±0.87
10.72 ±0.15
10.15 ±0.08
Table 3: Comparison of WikiText Perplexity ( ↓ ) . We evaluate Base pruning methods (Wanda and SparseGPT) and their variants augmented with FARMS, ATP, and SWA at sparsity 0.5/0.7. Results are averaged over three seeds and reported as mean ±std . The best result in each setting is bolded . SWA reduces perplexity and achieves the best values across all the models and sparsity levels.
Variant
SparseGPT
Wanda
0.5
0.7
0.5
0.7
SWA
12.10
20.10
11.90
52.56
SWA(Layer)
+3.00
+10.21
+0.08
+5.73
SWA(Head)
+3.69
+19.97
+0.17
+44.99
Table 4: Ablation study on hierarchical structural deletion. We report the absolute WikiText perplexity (PPL ↓ ) of SWA and the increase in PPL for each ablated variant relative to SWA on Qwen3 8B at sparsity levels of 0.5 and 0.7. SWA(Layer) uses only layer-level scores, and SWA(Head) uses only head-level scores.
Model
Base
GSM8K
ARC-C
MMLU
Llama3 8B
55.94 ±3.23
42.60 ±1.22
41.10 ±1.35
40.75 ±1.07
Qwen3 4B
50.93 ±1.62
40.42 ±1.12
38.90 ±1.23
39.89 ±2.35
Llama3.2 3B
70.70 ±5.07
50.54 ±1.30
49.69 ±2.69
48.95 ±1.20
Llama3.2 1B
157.20 ±3.20
109.2 ±6.20
110.9 ±2.60
105.28 ±6.62
Table 5: Effect of the graph-construction benchmark on SWA. We report WikiText perplexity ( ↓ ) for SparseGPT at 70% sparsity, comparing the base method with SWA using functional graphs built from GSM8K, ARC-C, or MMLU.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Training step
Tokens seen
Level 1
Level 2
NLL/token ↓
SWI ↑
NLL/token ↓
SWI ↑
4,000
8B
0.5089 ±0.16
0.6760 ±0.85
0.8255 ±0.25
0.4587 ±0.67
11,000
23B
0.4677 ±0.20
0.8008 ±0.13
0.7020 ±0.68
0.5574 ±0.80
30,000
63B
0.4344 ±0.28
0.8035 ±0.12
0.7007 ±0.52
0.5599 ±0.54
65,000
136B
0.2250 ±0.30
0.8135 ±0.19
0.6411 ±0.34
0.8100 ±0.32
71,000
149B
0.1717 ±0.07
0.8320 ±0.47
0.4026 ±0.18
0.8469 ±0.33
Appendix
Table 6: Relationship between DRE-Bench performance and small-world organization across Pythia-6.9B training checkpoints. We report the mean per-token negative log-likelihood over the answer tokens (NLL/token; lower is better) and SWI of the head-head functional graph on DRE-Bench Level 1 and Level 2 tasks. Results are averaged over three seeds and reported as mean ±std . Across training checkpoints, lower NLL is associated with higher SWI on both levels.
Model
Level 1
Level 2
NLL/token ↓
SWI ↑
NLL/token ↓
SWI ↑
Qwen3 14B
0.0822 ±0.01
1.8160 ±0.27
0.2410 ±0.03
1.7020 ±0.12
Qwen3 8B
0.0949 ±0.01
1.5850 ±0.28
0.2814 ±0.02
1.6000 ±0.07
Llama3 8B
0.0980 ±0.01
1.3750 ±0.19
0.2993 ±0.02
1.5935 ±0.11
Qwen3 4B
0.1142 ±0.02
1.2180 ±0.26
0.3397 ±0.05
1.2722 ±0.18
Llama3.2 3B
0.1377 ±0.07
1.1933 ±0.33
0.3412 ±0.10
1.2100 ±0.17
Appendix
Table 7: Relationship between DRE-Bench performance and small-world organization. We report the mean per-token negative log-likelihood over the answer tokens (NLL/token; lower is better) and SWI of the head-head functional graph on DRE-Bench Level 1 and Level 2 tasks. Results are averaged over three seeds and reported as mean ±std . Across models, lower NLL is associated with higher SWI.
Model
Size
Sparsity
Wanda
SparseGPT
Base
+ FARMS
+ ATP
+ SWA
Base
+ FARMS
+ ATP
+ SWA
Llama 3.2
1B
0.5
34.91 ±3.67
35.21 ±2.28
37.16 ±0.82
38.89 ±1.03
39.57 ±2.67
40.33 ±2.41
41.26 ±0.23
44.67 ±1.75
0.7
25.86 ±3.62
27.29 ±4.75
29.86 ±4.65
32.41 ±4.81
30.85 ±4.67
29.38 ±1.33
35.29 ±0.22
36.46 ±0.34
3B
0.5
40.53 ±1.41
42.71 ±1.22
44.57 ±1.39
45.80 ±1.87
45.50 ±0.36
46.67 ±0.50
46.05 ±1.41
49.91 ±0.31
0.7
40.16 ±2.41
41.50 ±1.63
40.91 ±4.74
43.04 ±0.87
40.42 ±1.16
42.25 ±0.85
41.69 ±4.46
43.57 ±1.11
Llama 3
8B
0.5
52.12 ±1.25
53.49 ±2.17
52.43 ±1.32
55.13 ±1.52
48.53 ±2.24
50.28 ±1.30
53.48 ±0.61
55.61 ±0.85
Appendix
Table 8: Comparison of Zero-shot Accuracy ( ↑ ). We evaluate Base pruning methods (Wanda and SparseGPT) and their variants augmented with FARMS, ATP, and SWA at sparsity 0.5/0.7. Results are averaged over three seeds and reported as mean ±std . The best result in each setting is bolded . SWA improves zero-shot accuracy and achieves the strongest performance over all the models and sparsity levels.
Method
Move
Change Color
Copy
Mirror
Fill Internal
Scale
Qwen3 8B (Non-Pruned)
36.00 ±4.11
23.14 ±4.61
12.18 ±4.12
84.66 ±3.20
6.07 ±1.99
83.33 ±5.87
SparseGPT
44.21 ±9.20
48.87 ±10.40
60.99 ±10.67
94.76 ±8.21
34.25 ±4.67
88.91 ±8.03
SparseGPT + FARMS
46.22 ±8.21
50.77 ±15.85
56.60 ±12.80
96.66 ±5.48
33.11 ±5.11
88.06 ±8.95
SparseGPT + ATP
43.09 ±9.01
47.73 ±8.66
56.79 ±12.48
93.56 ±9.63
32.66 ±5.81
89.49 ±7.47
SparseGPT + SWA
40.29 ±9.25
46.28 ±9.62
56.11 ±11.62
90.21 ±8.89
28.69 ±4.28
85.10 ±7.28
Appendix
Table 9: ARAOC task mismatch rate at sparsity 0.5. We report mismatch rate (Not M ↓ ) across six tasks. Results are averaged over three runs and reported as mean ±std . The best result in each setting is bolded .