When Can Attention Heads Be Statically Defined?
Organizations: University of Edinburgh, United Kingdom · Chamber of Deputies, Brazil
Abstract
Some attention heads learn similar patterns across inputs. Reusing these patterns could reduce training cost by avoiding repeated query-key score computation and softmax. Through controlled pretraining comparisons, we identify Selective Attention Freezing (SAF), which selects heads with low attention-pattern variance and replaces their attention weights with fitted post-softmax means halfway through training. We represent these fixed patterns with absolute-position and relative-distance preferences, reducing storage from quadratic to linear in sequence length. A fused kernel reconstructs the patterns and executes ordinary-attention and replaced heads together. At matched training-token budgets, replacing 25% of attention heads gives 1.056x faster post-replacement optimiser updates at 124M parameters and 4K context, with a 0.77% perplexity increase. At 1B and 8K context, post-replacement updates are 1.068x faster on four GPUs including communication, with a 0.51% perplexity increase. The resulting models also accelerate long-input finetuning and causal prefill. After associative-recall adaptation, the 124M model with 25% replacement generalises to more key-value pairs at a fixed length better than ordinary attention and two pruning controls.
Figures & tables
| Matched tokens: 2.4576B | Matched time | |||||
|---|---|---|---|---|---|---|
| (%) | Time (min) | (%) | ||||
| Operation + selection | 25% | 50% | 25% | 50% | 25% | 50% |
| Mean + variance | 184.7 | 178.7 | ||||
| Pruning + variance | 178.9 | 168.3 | ||||
| Mean + random heads | 183.4 | 178.8 | ||||
| (a) Pretraining on FineWeb-Edu | Perplexity | Update time | Memory | ||||
|---|---|---|---|---|---|---|---|
| Context ( ) | Replaced heads | Ordinary attn. | Replaced | PPL (%) | Ord./repl. (ms) | Speedup | Peak change |
| 4K (8) | 25% | 23.862 | 24.045 | 2249.49/2129.42 | 1.056 | ||
| 4K (8) | 50% | 23.862 | 24.457 | 2229.80/1991.85 | 1.119 | ||
| 8K (4) | 25% | 24.179 | 24.318 | 2710.24/2636.22 | 1.030 | ||
| 8K (4) | 50% | 24.179 | 24.676 | 2782.97/2524.39 | 1.101 | ||
| 16K (2) | 25% | 24.428 | 24.538 | 3741.42/3649.78 | 1.025 | ||
| Finetuning accuracy (%) | ||||
|---|---|---|---|---|
| Model | PPL | SST-2 | BoolQ | QuALITY |
| Ordinary attention | 12.446 | |||
| SAF 25% | 12.509 | |||
| Pruning 25% (gate-Taylor) | 12.529 | |||
| SAF 50% | 12.698 | |||
Appendix figures & tables31 assets
Supplementary material from the paper’s appendix.
Appendix
| Distance from mean | 10% | 25% | 50% | 75% | Stratified |
|---|---|---|---|---|---|
| Squared error | 0.900 | 1.000 | 1.000 | 0.900 | 0.950 |
| 0.600 | 0.600 | 0.600 | 0.700 | 0.625 | |
| Squared Hellinger | 0.800 | 0.900 | 0.900 | 1.000 | 0.900 |
| Total variation | 0.800 | 0.900 | 0.900 | 1.000 | 0.900 |
| Risk score | Spearman | 25% mean cost | 50% mean cost |
|---|---|---|---|
| Attention variance, | 0.723 | 0.058% | 0.115% |
| Residual cosine, | 0.810 | 0.033% | 0.121% |
| Relative attention-output change, | 0.703 | 0.052% | 0.150% |
| Projected-output magnitude, | 0.654 | 0.166% | 0.129% |
| Projected-output redundancy, | 0.320% | 0.367% |
| Context | Replaced | Variance | Forward KL | KL–variance |
|---|---|---|---|---|
| heads | PPL | PPL | (pp) | |
| 4K | 25% | 0.665% | 1.068% | +0.402 |
| 4K | 50% | 2.649% | 3.554% | +0.905 |
| 8K | 25% | 0.557% | 0.836% | +0.279 |
| 8K | 50% | 2.191% | 2.988% | +0.797 |
| 16K | 25% | 0.474% | 0.876% | +0.402 |
| Method choice | Alternatives evaluated |
|---|---|
| Fixed pattern | Post-softmax mean; sharp mean; Gaussian sample; Dirichlet sample; structured random |
| Head selection | Attention variance; row-wise forward KL; Q/K gradient norm; rank combinations; projected-output magnitude and redundancy; attention-residual influence |
| Amount and timing | Rates: 10%, 25%, 50%, 75%; times: 25%, 40%, 50%, 60%, 75%, 100%; one-time, gradual, validation-budget, and train-loss-triggered schedules |
| Representation and execution | Exact dense or absolute-plus-relative pattern; reference or fused mixed-head execution |
| (%) | ||
|---|---|---|
| 25% | 3750 | 0.724 |
| 40% | 3000 | 0.679 |
| 50% | 2500 | 0.662 |
| 60% | 2000 | 0.711 |
| 75% | 1250 | 0.936 |
| 100% | 0 | 8.168 |
| Comparison | Setting | Rate | (%) |
|---|---|---|---|
| Preset schedule | one replacement | 25.00% | 0.662 |
| gradual, five groups | 25.00% | 0.689 | |
| one replacement | 50.00% | 2.649 | |
| gradual, five groups | 50.00% | 2.832 | |
| Adaptive schedule | validation budget 1% | 6.25% | 0.093 |
| validation budget 2% | 12.50% | 0.348 |
| Rate | One-time (%) | Train-loss (%) | Difference (pp) | Time ratio (triggered/one-time) |
|---|---|---|---|---|
| 25% | 1.027 | |||
| 50% | 1.066 |
| Replaced heads | Update speedup | Peak allocated | Saving |
|---|---|---|---|
| 0% | 1.000 | 32.21 GiB | 0.00 GiB |
| 10% | 1.016 | 31.85 GiB | 0.36 GiB |
| 25% | 1.056 | 31.58 GiB | 0.63 GiB |
| 50% | 1.119 | 31.24 GiB | 0.98 GiB |
| 75% | 1.215 | 30.84 GiB | 1.38 GiB |
| Paired differences | |||||
|---|---|---|---|---|---|
| Setting (rate) | PPL | Training (min) | Intervention (s) | (pp) | Time (s) |
| Ordinary attention | — | — | — | ||
| Mean + variance (25%) | — | — | |||
| Mean + variance (50%) | — | — | |||
| Pruning + variance (25%) | |||||
| Pruning + variance (50%) | |||||
| Model | Rate | Capture and mean | Fit | Total intervention |
|---|---|---|---|---|
| 124M, 4K | 25% | |||
| 124M, 4K | 50% | |||
| 1B, 8K | 25% | 1422.37 | 39.56 | 1464.41 |
| 1B, 8K | 50% | 1313.35 | 76.28 | 1392.64 |
| (a) Matched-token pretraining: 124M, 4K, 2.4576B tokens | ||||
|---|---|---|---|---|
| Pattern | PPL | Update (s) | Continuation (min) | Peak (GiB) |
| Frozen mean | ||||
| Trainable mean | ||||
| Operation + selection | Rate | PPL | Updates | Time (min) |
|---|---|---|---|---|
| Ordinary attention | 0% | |||
| Mean + variance | 25% | |||
| Mean + variance | 50% | |||
| Pruning + variance | 25% | |||
| Pruning + variance | 50% | |||
| Mean + random heads | 25% |
| Pretraining PPL | Finetuning accuracy (%) | |||
|---|---|---|---|---|
| Operation + selection | Matched tokens | Matched time | SST-2 | BoolQ |
| Ordinary attention | ||||
| SAF + variance | ||||
| Pruning + variance | ||||
| Pruning + gate-Taylor | ||||
| Accuracy (%) | |||
| Operation + selection | Rate | SST-2 | BoolQ |
| Midpoint, matched time: three pretraining seeds, one finetuning seed | |||
| Ordinary attention | 0% | ||
| Mean + variance | 25% | ||
| Mean + variance | 50% | ||
| Pruning + variance | 25% | ||
| Pruning 25% | |||||
| Updates | Tokens / pairs | Ordinary attn. | SAF 25% | Variance | Gate-Taylor |
| Adaptation: three pretraining seeds, task seed 1337 | |||||
| 1,500 | 512 / 8 | ||||
| 512 / 16 | |||||
| 512 / 24 | |||||
| 512 / 32 | |||||
| Update time (ms) | Peak change (%) | ||||||
|---|---|---|---|---|---|---|---|
| Context | Rate | Ordinary attn. | Mean | Pruning | Mean | Pruning | |
| 4K | 20 | 25% | 2172.80 | 2057.35 | 1923.26 | ||
| 4K | 20 | 50% | 2173.12 | 1929.10 | 1678.24 | ||
| 8K | 10 | 25% | 2651.80 | 2573.68 | 2307.18 | ||
| 8K | 10 | 50% | 2652.83 | 2408.64 | 1933.85 | ||
| 16K | 5 | 25% | 3714.63 | 3621.11 | 3083.03 | ||
| Operation + selection | Ordinary attn. (ms) | Method (ms) | Speedup ( ) | Peak change (%) |
| Matched-token pretraining | ||||
| Mean + variance | 154.59 | 142.52 | ||
| Pruning + variance | 155.26 | 133.43 | ||
| Pruning + gate-Taylor | 154.22 | 132.63 | ||
| Pruning from initialisation | 154.27 | 132.64 | ||
| Matched-time pretraining | ||||
| Task / checkpoint | Accuracy (%) | accuracy (pp) | Time/update | Speedup | Peak memory change | |
|---|---|---|---|---|---|---|
| SST-2, ordinary attn. | 256 | — | ms | 1.000 | 0.00% | |
| SST-2, fixed 25% | 256 | ms | ||||
| SST-2, fixed 50% | 256 | ms | ||||
| BoolQ, ordinary attn. | 64 | — | ms | 1.000 | 0.00% | |
| BoolQ, fixed 25% | 64 | ms | ||||
| BoolQ, fixed 50% | 64 | ms |
| Task | Schedule | accuracy (pp) | Time/update | Controller cost |
|---|---|---|---|---|
| SST-2 | One-time | 33.27 ms | 0.70 s | |
| Train-loss | 36.15 ms | 45.12 s | ||
| Post-training | 30.75 ms | 0.71 s | ||
| BoolQ | One-time | 40.85 ms | 6.18 s | |
| Train-loss | 66.34 ms | 70.15 s | ||
| Post-training | 31.44 ms | 6.17 s |
| GPUs / | Rate | Ordinary attn. (ms) | SAF (ms) | Speedup ( ) | Peak change (%) |
| Four GPUs, including communication: 524,288 tokens per update | |||||
| 4 / 2 | 25% | 4258.67 | 3986.67 | — | |
| 4 / 2 | 50% | 4263.01 | 3653.78 | — | |
| Single GPU, no communication: 131,072 tokens per update | |||||
| 1 / 1 | 25% | 4224.85 | 3958.26 | 1.067 | |
| 1 / 1 | 50% | 4226.87 | 3632.25 | 1.163 | |
| Accuracy by seed | ||||||
|---|---|---|---|---|---|---|
| Task | Model | 1337 | 1338 | 1339 | Change (pp) | Update (ms) |
| SST-2 | Ordinary attention | 91.63 | 90.94 | 91.74 | — | |
| SAF 25% | 90.94 | 91.86 | 92.09 | |||
| Pruning 25% | 91.06 | 91.86 | 92.20 | |||
| SAF 50% | 91.28 | 91.51 | 90.14 | |||
| BoolQ | Ordinary attention | 75.99 | 76.39 | 75.57 | — | |
| Attention variance | Forward KL | ||||||
| Task | Ordinary attn. | 10% | 20% | 30% | 10% | 20% | 30% |
| HellaSwag | 52.28 | 52.02 | 51.21 | 48.92 | 49.10 | 45.05 | 43.81 |
| PIQA | 74.86 | 75.19 | 74.48 | 73.56 | 74.70 | 72.58 | 70.40 |
| ARC-Easy | 80.43 | 79.63 | 77.61 | 75.72 | 72.90 | 68.06 | 64.10 |
| SST-2 | 88.65 | 88.65 | 87.96 | 85.44 | 87.61 | 63.07 | 68.81 |
| BoolQ | 84.46 | 83.33 | 80.49 | 73.12 | 80.03 | 71.38 | 71.38 |
| 10% | 20% | 30% | 50% | 75% | ||
|---|---|---|---|---|---|---|
| 4 | 2048 | 1.011 (1.16) | 1.028 (1.18) | 1.068 (1.21) | 1.135 (1.29) | 1.225 (1.38) |
| 4 | 4096 | 1.027 (1.16) | 1.057 (1.19) | 1.097 (1.23) | 1.156 (1.28) | 1.246 (1.38) |
| 4 | 8192 | 1.042 (1.15) | 1.072 (1.19) | 1.115 (1.23) | 1.184 (1.30) | 1.250 (1.37) |
| 8 | 2048 | 1.015 (1.15) | 1.042 (1.17) | 1.079 (1.21) | 1.148 (1.28) | 1.231 (1.37) |
| 8 | 4096 | 1.029 (1.15) | 1.061 (1.19) | 1.080 (1.18) | 1.156 (1.27) | 1.252 (1.37) |
| 8 | 8192 | 1.040 (1.15) | 1.071 (1.18) | 1.110 (1.22) | 1.178 (1.30) | 1.272 (1.40) |
| Position | Ordinary attn. PPL | Dense PPL | Abs.+rel. PPL | Storage reduction |
|---|---|---|---|---|
| Absolute | 26.0252 | 26.1184 | 26.1150 | 59.9 |
| RoPE | 23.9571 | 24.1180 | 24.1165 | 60.6 |
| FP32 | BF16 autocast | |||
|---|---|---|---|---|
| Tensor | Max. absolute | Relative | Max. absolute | Relative |
| Output | ||||
| Input gradient | ||||
| Value gradient | ||||
| Query projection | ||||
| Key projection | ||||
| Input length | Batch size | Replaced heads | Ordinary attn. (ms) | Replaced (ms) | Speedup |
|---|---|---|---|---|---|
| 4K | 64 | 25% | 306.70 | 278.33 | 1.100 |
| 4K | 64 | 50% | 303.94 | 247.47 | 1.229 |
| 8K | 64 | 25% | 776.17 | 705.03 | 1.100 |
| 8K | 64 | 50% | 760.20 | 622.94 | 1.221 |
| 16K | 64 | 25% | 2075.04 | 1900.13 | 1.092 |
| 16K | 64 | 50% | 2045.38 | 1703.87 | 1.200 |
| Asset | Use | Upstream licence or terms |
| Datasets | ||
| FineWeb-Edu | Pretraining; calibration and evaluation | ODC-By 1.0 ; Common Crawl terms |
| SST-2 (GLUE) | Finetuning and zero-shot evaluation | Not specified in the original dataset card |
| BoolQ (SuperGLUE) | Finetuning and zero-shot evaluation | CC BY-SA 3.0 |
| QuALITY | Long-input finetuning | CC BY 4.0 ; article-level licences |
| HellaSwag | Zero-shot evaluation | MIT |