Effective Does Not Mean Useful: Conditional Functional Substitutability for Redundancy and Scaling in Transformers
Authors: Jiaheng Chen, Jiaxing Li, Yucheng Xiao, Xinyong Cai, Juncheng Bu, Lan Yu, Tinghe Zhang
Organizations: Harbin Institute of Technology, Shenzhen, China · Northeastern University, Shenyang, China · Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China · Peking University Shenzhen Graduate School, Shenzhen, China · Tsinghua University, Beijing, China
Modern neural networks scale predictably, yet the mechanisms behind these regularities remain unclear. Neural redundancy is typically characterized by component importance or representational similarity, both indirect proxies. We view redundancy as an input-conditioned, dynamic relation: intermediate computational states are functionally redundant when they induce similar downstream responses. We introduce Conditional Functional Substitutability (CFS) to directly characterize such functional substitution. CFS exposes functional relations and reduction potential missed by conventional importance- and similarity-based measures. Across modalities and Transformer families, CFS reveals systematic functional reorganization with scale. Controlled scaling further shows that performance gains need not track growth in substitutability, while fixed-capacity models with more independent functional structure perform better, providing a functional account of diminishing returns. Predicted CFS further enables dynamic computation with a better performance--computation trade-off than importance-based component selection, suggesting new directions for redundancy-aware computation and more efficient model scaling.
Figures & tables
Figure 1: A locally effective computation may lack independent global utility when its downstream role is absorbed by an alternative path. CFS directly characterizes this functional substitutability.
Exact Oracle
Taylor
Keep
KL ↓
Acc.
Fid.
KL ↓
Acc.
Fid.
3/12
0.08119
81.47
96.88
0.21156
80.58
93.75
6/12
0.01290
82.37
99.55
0.05125
81.47
95.76
9/12
0.00258
81.70
99.55
0.01375
81.03
98.44
Table 2: Exact joint-subset oracle versus Taylor on ViT-B/16. Fidelity is agreement with the dense prediction; dense accuracy is 81.92% .
Family
Scale
Params
Metric
Raw
Match-2
Cover/H
Rank/H
ViT
Tiny
5.72M
76.95
0.3379
0.3379
0.7982
0.8998
Small
22.05M
83.40
0.4088
0.2852
0.7321
0.8602
Base
86.57M
85.35
0.4976
0.2889
0.5972
0.7276
Large
304.33M
86.91
0.6034
0.3737
0.3455
0.4647
BERT
Tiny
4.42M
4.594
0.0128
0.0128
1.0000
0.9998
Small
28.80M
3.033
0.1318
0.0634
0.9922
0.9827
Table 3: Functional organization across natural scaling families. Metric is accuracy for ViT and LM loss for BERT/Qwen2.5; Cover/H and Rank/H are normalized. Qwen2.5-0.5B is omitted due to its 64-d rather than 128-d query heads.
Axis
Setting
Params
Val. Loss
Raw
Match-2
Cover/H
Rank/H
Width
W576
77.40M
3.7021
0.1882
0.1140
0.9826
0.9550
W768
124.44M
3.5752
0.2125
0.1241
0.9766
0.9335
W1152
250.36M
3.4000
0.2522
0.1325
0.9497
0.8994
Depth
D9
103.18M
3.6252
0.2406
0.1264
0.9236
0.9159
D12
124.44M
3.5752
0.2125
0.1241
0.9766
0.9335
D18
166.97M
3.5024
0.2154
0.1334
0.9722
0.9286
Table 4: Controlled width/depth scaling under a shared training protocol. W768 and D12 are the same model.
Heads
Params
Val. Loss
Raw
Match-2
Cover/H
Rank/H
Eff. Rank
6×128
124.44M
3.5707
0.2012
0.1413
0.9497
0.9507
5.704
12×64
124.44M
3.5752
0.2125
0.1241
0.9766
0.9335
11.202
24×32
124.44M
3.5780
0.2583
0.1371
0.9518
0.8725
20.939
Table 5: Functional organization at fixed model capacity. All models have 12 layers, width 768, and identical parameter counts; only head decomposition changes.
Figure 2: Formation of functional organization during training. Left: CFS-graph similarity (top) and top-1 substitute agreement (bottom) with epoch 100. Right: per-head similarity to final substitute rankings for layers 3, 7, and 11.
Information
Recovery
Rep. Acc.
Spearman
Residual
0.2204
36.85
0.2553
Q/K moments
0.2251
38.82
0.2718
Attention logits
0.2232
38.24
0.2678
V content
0.2245
38.65
0.2790
AV
0.2236
38.53
0.2732
Post- WO
0.2247
39.03
0.2793
Table 6: CFS observability on DeiT-S. Recovery is achieved by the predicted representative.
Keep
CFS Acc.
Taylor Acc.
CFS KL
Taylor KL
KL Red.
3/12
81.03
80.58
0.1190
0.2105
43.46%
6/12
82.59
81.47
0.0281
0.0513
45.21%
9/12
82.07
81.03
0.0070
0.0138
49.23%
Table 7: CFS routing versus Taylor under equal logical head budgets. KL is measured against the dense model.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Performance and functional organization throughout training. Validation performance improves steadily, while functional substitutability, coverage, and effective rank follow distinct trajectories. Predictive learning and functional reorganization are therefore not synchronized.
Figure 4: Depth-resolved functional consolidation. Effective functional rank and functional coverage evolve differently across layers 3, 7, and 11, indicating that functional consolidation proceeds at different rates and to different degrees throughout the network.
Figure 5: Evolution of the mean CFS graph during training. We visualize layers 3, 7, and 11 at epochs 1, 25, 50, 75, and 100. For each source head, only its strongest positive outgoing substitution edge is shown; node size reflects mean deletion sensitivity. These snapshots provide a qualitative view of the progressive reorganization quantified in Section 5.4 .
Setting
CFS samples
Evaluated layers
α search
Valid source
Coverage
ViT natural
512
final block
31-point [0,3]
Ddrop>10−5
τ=0.5
BERT natural
32
final layer
16-point grid
Ddrop>10−5
τ=0.5
Qwen2.5 natural
32
≈25/50/75% depth
16-point grid
Ddrop>10−5
τ=0.5
Training dynamics
512
layers 3/7/11
16-point [0,3]
Ddrop>10−5
τ=0.25/0.5/0.75
Appendix
Table 8: CFS protocol for the natural-family and training-dynamics experiments. Matched-2 denotes the expected best substitute among two candidates.
A central question in modern machine learning is how much a trained model can be compressed without changing its behavior, to reduce the memory, compute and energy required to deploy it. To study this, we quantify functional degeneracy through the behavioral recovery rank, defined as the number of leading behavioral-Hessian eigendirections required to recover a trained model's performance. Using the behavioral recovery rank as a geometric benchmark for compression, we find that structural and magnitude pruning retain more degrees of freedom, even after the task is saturated. This gap suggests that functional redundancy is distributed across parameter directions and is not exposed by individual weights or neurons.
Maria Matveev, Pascal Esser, Ayush Bharadwaj +2
Department of Mathematics, LMU Munich · Munich Center for Machine Learning (MCML) · Independent +3
Mechanistic interpretability seeks to explain transformer behavior through circuits: sets of internal components that causally support a behavior. However, self-repair creates a blind spot: ablating a primary component can activate a dormant backup, so a circuit that explains behavior in the intact model can become incomplete under the intervention used to test it. We formulate this gap as conditional circuit completion: given a primary set, identify components that become causally important after its removal. We introduce conditional co-ablation (CoAx), which ranks candidates by growth in ablation effect after primary-set removal. We show that a perfectly dormant backup can be indistinguishable from an irrelevant component to per-unit intact-state scores, whereas its conditional effect change exactly aggregates all interaction orders linking it to the removed set. On GPT-2-small's Indirect Object Identification (IOI) circuit, CoAx recovers the documented backup heads at 0.941 ROC-AUC, versus 0.815 for the strongest intact-state attribution baseline and 0.758 for the matched conditional-energy control. Recovery drops to 0.40 +/- 0.13 AUC for alternative component sets matched in behavioral effect, output displacement, and depth, showing that recovery is specific to the removed circuit. Beyond recovery, the CoAx-selected heads are causally load-bearing: freezing them after primary removal sharply reduces the IOI margin, while adding them to the incomplete circuit reduces incompleteness from 0.75 to 0.21. More broadly, conditional growth aligns with intervention-derived repair in 11/12 held-out instances across 4 mechanism clusters, and CoAx completions outperform matched random completions on all 8 non-GPT-2 models spanning 6 architecture families. Together, causal explanations of self-repairing transformers must account for backup circuitry when primary components fail.
Zhiren Gong, He Lu, Tiantong Wang +6
College of Computing and Data Science, Nanyang Technological University, Singapore · Interdisciplinary Graduate Programme, Nanyang Technological University, Singapore · School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore
When researchers ask whether two transformer layers are "equivalent" for compression, they often conflate distinct tests. Replacement asks whether one layer's map can substitute for another's in place; interchange asks whether two layers approximately commute when their positions are swapped. Both are output-grounded swap-KL probes, but they need not agree: on pretrained transformers the protocol gap can change which layers look safe to prune by several-fold under the same evaluator, especially when replacement distances are high. We measure both protocols across checkpoints and architectures. On a Pythia training trajectory (410M and 1.4B), the replacement-interchange gap grows from initialization to convergence. Under one matched WikiText-2 contract at 8B scale, Qwen3-8B enters a divergent regime: interchange-guided removal is several-fold safer than replacement-guided at the same layer budgets, while Llama-3.1-8B ties the two protocols for pruning cost even though interchange KL is lower, showing metric gaps need not map one-to-one to removal. Before layer removal or merging, score both swap-KLs on the target checkpoint; the diagnostic requires only unlabeled forward passes.