Effective Does Not Mean Useful: Conditional Functional Substitutability for Redundancy and Scaling in Transformers
Authors: Jiaheng Chen, Jiaxing Li, Yucheng Xiao, Xinyong Cai, Juncheng Bu, Lan Yu, Tinghe Zhang
Organizations: Harbin Institute of Technology, Shenzhen, China · Northeastern University, Shenyang, China · Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China · Peking University Shenzhen Graduate School, Shenzhen, China · Tsinghua University, Beijing, China
Modern neural networks scale predictably, yet the mechanisms behind these regularities remain unclear. Neural redundancy is typically characterized by component importance or representational similarity, both indirect proxies. We view redundancy as an input-conditioned, dynamic relation: intermediate computational states are functionally redundant when they induce similar downstream responses. We introduce Conditional Functional Substitutability (CFS) to directly characterize such functional substitution. CFS exposes functional relations and reduction potential missed by conventional importance- and similarity-based measures. Across modalities and Transformer families, CFS reveals systematic functional reorganization with scale. Controlled scaling further shows that performance gains need not track growth in substitutability, while fixed-capacity models with more independent functional structure perform better, providing a functional account of diminishing returns. Predicted CFS further enables dynamic computation with a better performance--computation trade-off than importance-based component selection, suggesting new directions for redundancy-aware computation and more efficient model scaling.
Figures & tables
Figure 1: A locally effective computation may lack independent global utility when its downstream role is absorbed by an alternative path. CFS directly characterizes this functional substitutability.
Exact Oracle
Taylor
Keep
KL ↓
Acc.
Fid.
KL ↓
Acc.
Fid.
3/12
0.08119
81.47
96.88
0.21156
80.58
93.75
6/12
0.01290
82.37
99.55
0.05125
81.47
95.76
9/12
0.00258
81.70
99.55
0.01375
81.03
98.44
Table 2: Exact joint-subset oracle versus Taylor on ViT-B/16. Fidelity is agreement with the dense prediction; dense accuracy is 81.92% .
Family
Scale
Params
Metric
Raw
Match-2
Cover/H
Rank/H
ViT
Tiny
5.72M
76.95
0.3379
0.3379
0.7982
0.8998
Small
22.05M
83.40
0.4088
0.2852
0.7321
0.8602
Base
86.57M
85.35
0.4976
0.2889
0.5972
0.7276
Large
304.33M
86.91
0.6034
0.3737
0.3455
0.4647
BERT
Tiny
4.42M
4.594
0.0128
0.0128
1.0000
0.9998
Small
28.80M
3.033
0.1318
0.0634
0.9922
0.9827
Table 3: Functional organization across natural scaling families. Metric is accuracy for ViT and LM loss for BERT/Qwen2.5; Cover/H and Rank/H are normalized. Qwen2.5-0.5B is omitted due to its 64-d rather than 128-d query heads.
Axis
Setting
Params
Val. Loss
Raw
Match-2
Cover/H
Rank/H
Width
W576
77.40M
3.7021
0.1882
0.1140
0.9826
0.9550
W768
124.44M
3.5752
0.2125
0.1241
0.9766
0.9335
W1152
250.36M
3.4000
0.2522
0.1325
0.9497
0.8994
Depth
D9
103.18M
3.6252
0.2406
0.1264
0.9236
0.9159
D12
124.44M
3.5752
0.2125
0.1241
0.9766
0.9335
D18
166.97M
3.5024
0.2154
0.1334
0.9722
0.9286
Table 4: Controlled width/depth scaling under a shared training protocol. W768 and D12 are the same model.
Heads
Params
Val. Loss
Raw
Match-2
Cover/H
Rank/H
Eff. Rank
6×128
124.44M
3.5707
0.2012
0.1413
0.9497
0.9507
5.704
12×64
124.44M
3.5752
0.2125
0.1241
0.9766
0.9335
11.202
24×32
124.44M
3.5780
0.2583
0.1371
0.9518
0.8725
20.939
Table 5: Functional organization at fixed model capacity. All models have 12 layers, width 768, and identical parameter counts; only head decomposition changes.
Figure 2: Formation of functional organization during training. Left: CFS-graph similarity (top) and top-1 substitute agreement (bottom) with epoch 100. Right: per-head similarity to final substitute rankings for layers 3, 7, and 11.
Information
Recovery
Rep. Acc.
Spearman
Residual
0.2204
36.85
0.2553
Q/K moments
0.2251
38.82
0.2718
Attention logits
0.2232
38.24
0.2678
V content
0.2245
38.65
0.2790
AV
0.2236
38.53
0.2732
Post- WO
0.2247
39.03
0.2793
Table 6: CFS observability on DeiT-S. Recovery is achieved by the predicted representative.
Keep
CFS Acc.
Taylor Acc.
CFS KL
Taylor KL
KL Red.
3/12
81.03
80.58
0.1190
0.2105
43.46%
6/12
82.59
81.47
0.0281
0.0513
45.21%
9/12
82.07
81.03
0.0070
0.0138
49.23%
Table 7: CFS routing versus Taylor under equal logical head budgets. KL is measured against the dense model.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Performance and functional organization throughout training. Validation performance improves steadily, while functional substitutability, coverage, and effective rank follow distinct trajectories. Predictive learning and functional reorganization are therefore not synchronized.
Figure 4: Depth-resolved functional consolidation. Effective functional rank and functional coverage evolve differently across layers 3, 7, and 11, indicating that functional consolidation proceeds at different rates and to different degrees throughout the network.
Figure 5: Evolution of the mean CFS graph during training. We visualize layers 3, 7, and 11 at epochs 1, 25, 50, 75, and 100. For each source head, only its strongest positive outgoing substitution edge is shown; node size reflects mean deletion sensitivity. These snapshots provide a qualitative view of the progressive reorganization quantified in Section 5.4 .
Setting
CFS samples
Evaluated layers
α search
Valid source
Coverage
ViT natural
512
final block
31-point [0,3]
Ddrop>10−5
τ=0.5
BERT natural
32
final layer
16-point grid
Ddrop>10−5
τ=0.5
Qwen2.5 natural
32
≈25/50/75% depth
16-point grid
Ddrop>10−5
τ=0.5
Training dynamics
512
layers 3/7/11
16-point [0,3]
Ddrop>10−5
τ=0.25/0.5/0.75
Appendix
Table 8: CFS protocol for the natural-family and training-dynamics experiments. Matched-2 denotes the expected best substitute among two candidates.
College of Computing and Data Science, Nanyang Technological University, Singapore · Interdisciplinary Graduate Programme, Nanyang Technological University, Singapore · School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore