Some attention heads learn similar patterns across inputs. Reusing these patterns could reduce training cost by avoiding repeated query-key score computation and softmax. Through controlled pretraining comparisons, we identify Selective Attention Freezing (SAF), which selects heads with low attention-pattern variance and replaces their attention weights with fitted post-softmax means halfway through training. We represent these fixed patterns with absolute-position and relative-distance preferences, reducing storage from quadratic to linear in sequence length. A fused kernel reconstructs the patterns and executes ordinary-attention and replaced heads together. At matched training-token budgets, replacing 25% of attention heads gives 1.056x faster post-replacement optimiser updates at 124M parameters and 4K context, with a 0.77% perplexity increase. At 1B and 8K context, post-replacement updates are 1.068x faster on four GPUs including communication, with a 0.51% perplexity increase. The resulting models also accelerate long-input finetuning and causal prefill. After associative-recall adaptation, the 124M model with 25% replacement generalises to more key-value pairs at a fixed length better than ordinary attention and two pruning controls.
Figures & tables
Figure 1: Overview of fixed-pattern replacement of ordinary attention.
Figure 2: Compact representation and execution of a fixed attention pattern.
Figure 3: Perplexity increases for five fixed patterns and four replacement rates.
Matched tokens: 2.4576B
Matched time
ΔPPL (%)
Time (min)
ΔPPL (%)
Operation + selection
25%
50%
25%
50%
25%
50%
Mean + variance
0.663±0.018
2.495±0.024
184.7
178.7
0.456±0.108
1.911±0.112
Pruning + variance
0.806±0.018
3.033±0.056
178.9
168.3
0.113±0.169
0.978±0.108
Mean + random heads
1.282±0.056
3.630±0.093
183.4
178.8
1.243±0.112
2.778±0.125
Table 1: Perplexity increase (%, lower is better) at matched tokens or time: mean ± SE over three seeds. Times include the initial training with ordinary attention and intervention.
(a) Pretraining on FineWeb-Edu
Perplexity ↓
Update time
Memory
Context ( B )
Replaced heads
Ordinary attn.
Replaced
Δ PPL (%)
Ord./repl. (ms)
Speedup
Peak change
4K (8)
25%
23.862
24.045
+0.768±0.067%
2249.49/2129.42
1.056 ×
−1.96%
4K (8)
50%
23.862
24.457
+2.492±0.083%
2229.80/1991.85
1.119 ×
−3.04%
8K (4)
25%
24.179
24.318
+0.578±0.011%
2710.24/2636.22
1.030 ×
−1.63%
8K (4)
50%
24.179
24.676
+2.058±0.066%
2782.97/2524.39
1.101 ×
−2.55%
16K (2)
25%
24.428
24.538
+0.450±0.012%
3741.42/3649.78
1.025 ×
−0.82%
Table 2: Quality and training costs for ordinary-attention and SAF 124M checkpoints. Ord./repl. (ms) reports optimiser-update times for the ordinary-attention and replaced models.
Figure 4: SAF speedup over FlashAttention: (a) pretraining updates; (b) finetuning updates; (c) causal prefill by input length at B=64 ; (d) prefill by batch size at T=4096 .
Figure 5: MQAR accuracy for later-intervention, matched-time 124M checkpoints. Models train on eight pairs at 512 tokens. (The ordinary-attention and gate-Taylor pruning curves nearly overlap.)
Figure 6: Qwen3-4B zero-shot accuracy changes (pp) relative to the ordinary-attention model.
Finetuning accuracy (%) ↑
Model
PPL ↓
SST-2
BoolQ
QuALITY
Ordinary attention
12.446
91.44±0.25
75.98±0.24
27.44±1.25
SAF 25%
12.509
91.63±0.35
75.27±1.45
27.60±0.21
Pruning 25% (gate-Taylor)
12.529
91.70±0.34
75.46±1.19
27.66±0.71
SAF 50%
12.698
90.98±0.43
75.68±0.31
25.12±0.75
Table 3: 1B quality at 8K context and 19.667B training tokens. PPL uses one pretraining seed; task accuracies (%) show mean ± SE over three finetuning seeds.
Appendix figures & tables31 assets
Supplementary material from the paper’s appendix.
Appendix
Distance from mean
10%
25%
50%
75%
Stratified
Squared error
0.900
1.000
1.000
0.900
0.950
DKL(A∥P)
0.600
0.600
0.600
0.700
0.625
Squared Hellinger
0.800
0.900
0.900
1.000
0.900
Total variation
0.800
0.900
0.900
1.000
0.900
Appendix
Table 4: Fixed-pattern geometry: within-rate correlation between distance from the post-softmax mean and final perplexity increase.
Risk score
Spearman ↑
25% mean cost ↓
50% mean cost ↓
Attention variance, svar
0.723
0.058%
0.115%
Residual cosine, scos
0.810
0.033%
0.121%
Relative attention-output change, srel
0.703
0.052%
0.150%
Projected-output magnitude, smag
0.654
0.166%
0.129%
Projected-output redundancy, sred
−0.013
0.320%
0.367%
Appendix
Table 5: Immediate one-head diagnostics at the 4K RoPE midpoint. Spearman correlates scores with NLL changes; mean costs are relative perplexity increases.
Context
Replaced
Variance
Forward KL
KL–variance
heads
Δ PPL
Δ PPL
(pp)
4K
25%
0.665%
1.068%
+0.402
4K
50%
2.649%
3.554%
+0.905
8K
25%
0.557%
0.836%
+0.279
8K
50%
2.191%
2.988%
+0.797
16K
25%
0.474%
0.876%
+0.402
Appendix
Table 6: Final perplexity increase after joint head selection, midpoint replacement, and 2,500 continuation updates.
Figure 7: Selected-head diagnostics: (a) layer distribution; (b) quarter-to-midpoint retention; (c) midpoint attention profiles; (d) entropy and variance in the original one-head probe. Bars in (a,c) show mean ± SE over three seeds.
Figure 8: Paired train-loss gap after replacing 25% of heads at training fraction τ .
Method choice
Alternatives evaluated
Fixed pattern
Post-softmax mean; sharp mean; Gaussian sample; Dirichlet sample; structured random
Exact dense or absolute-plus-relative pattern; reference or fused mixed-head execution
Appendix
Table 7: Method choices evaluated in controlled 124M pretraining.
τ
M−t
ΔPPL (%)
25%
3750
0.724
40%
3000
0.679
50%
2500
0.662
60%
2000
0.711
75%
1250
0.936
100%
0
8.168
Appendix
Table 8: Timing at 25% replacement with the post-softmax mean. τ : fraction of training completed; M−t : remaining updates; ΔPPL : final perplexity increase.
Comparison
Setting
Rate
ΔPPL (%)
Preset schedule
one replacement
25.00%
0.662
gradual, five groups
25.00%
0.689
one replacement
50.00%
2.649
gradual, five groups
50.00%
2.832
Adaptive schedule
validation budget 1%
6.25%
0.093
validation budget 2%
12.50%
0.348
Appendix
Table 9: Single-seed comparison of preset and adaptive replacement schedules.
Rate
One-time ΔPPL (%)
Train-loss ΔPPL (%)
Difference (pp)
Time ratio (triggered/one-time)
25%
0.768±0.067
0.858±0.180
+0.091±0.123
1.027 ×
50%
2.492±0.083
2.495±0.189
+0.003±0.267
1.066 ×
Appendix
Table 10: Three-seed comparison of one-time and train-loss-triggered schedules at matched final rates.
Replaced heads
Update speedup
Peak allocated
Saving
0%
1.000 ×
32.21 GiB
0.00 GiB
10%
1.016 ×
31.85 GiB
0.36 GiB
25%
1.056 ×
31.58 GiB
0.63 GiB
50%
1.119 ×
31.24 GiB
0.98 GiB
75%
1.215 ×
30.84 GiB
1.38 GiB
Appendix
Table 11: Post-replacement training speed and memory at T=4096 and batch size eight.
Figure 9: Fixed-entry CDFs at i=1024 , j=1023 : logits (top) and attention probabilities (bottom, logarithmic horizontal axis).
Figure 10: UMAP of empirical binned attention profiles and fitted Dirichlet samples at concentration scale 4.
Paired differences
Setting (rate)
PPL ↓
Training (min)
Intervention (s)
ΔPPL (pp)
Time (s)
Ordinary attention
23.896±0.038
188.7±2.7
—
—
—
Mean + variance (25%)
24.054±0.041
184.7±1.6
54.9±6.0
—
—
Mean + variance (50%)
24.492±0.035
178.7±2.0
46.3±1.2
—
—
Pruning + variance (25%)
24.088±0.042
178.9±1.6
46.3±3.1
+0.1428±0.0040
−343.9±27.8
Pruning + variance (50%)
24.621±0.049
168.3±1.3
50.8±5.8
+0.5378±0.0622
−620.9±65.3
Appendix
Table 12: Matched-token controls: mean ± SE over three pretraining seeds. Paired columns subtract midpoint mean + variance at the same rate.
Model
Rate
Capture and mean
Fit α,ρ
Total intervention
124M, 4K
25%
51.56±5.96
3.20±0.16
54.86±6.04
124M, 4K
50%
40.89±1.11
5.24±0.09
46.30±1.22
1B, 8K
25%
1422.37
39.56
1464.41
1B, 8K
50%
1313.35
76.28
1392.64
Appendix
Table 13: One-time calibration and fitting costs in seconds. The 124M controls report mean ± SE over three seeds; the 1B runs use one seed. Fitting is included in the total.
Table 14: Pattern controls at 25% replacement: mean ± SE over three pretraining seeds.
Operation + selection
Rate
PPL ↓
Updates
Time (min)
Ordinary attention
0%
23.882±0.023
5008±13
187.65±0.25
Mean + variance
25%
23.991±0.006
5109±11
187.64±0.25
Mean + variance
50%
24.338±0.020
5221±18
187.64±0.25
Pruning + variance
25%
23.909±0.024
5267±30
187.64±0.25
Pruning + variance
50%
24.115±0.003
5685±36
187.65±0.25
Mean + random heads
25%
24.178±0.024
5097±28
187.64±0.25
Appendix
Table 15: Matched-time controls: mean ± SE over three pretraining seeds. Updates and times include the initial training with ordinary attention.
Pretraining PPL ↓
Finetuning accuracy (%) ↑
Operation + selection
Matched tokens
Matched time
SST-2
BoolQ
Ordinary attention
22.563±0.010
22.561±0.011
87.00±0.33
71.78±0.46
SAF + variance
22.923±0.013
22.895±0.012
87.73±0.46
71.27±0.67
Pruning + variance
23.019±0.028
22.953±0.026
87.81±0.34
71.75±0.29
Pruning + gate-Taylor
22.865±0.019
22.798±0.021
86.96±0.49
71.33±0.16
Appendix
Table 16: Later 25% intervention: mean ± SE across three pretraining seeds. Finetuning uses the matched-time checkpoints and one fixed seed.
Accuracy (%) ↑
Operation + selection
Rate
SST-2
BoolQ
Midpoint, matched time: three pretraining seeds, one finetuning seed
Ordinary attention
0%
86.24±0.93
71.86±0.40
Mean + variance
25%
86.05±0.54
70.52±0.16
Mean + variance
50%
86.12±0.54
70.66±0.26
Pruning + variance
25%
86.85±0.56
70.97±0.66
Appendix
Table 17: Supplementary finetuning comparisons. Each group specifies which seed varies; errors are standard errors.
Pruning 25%
Updates
Tokens / pairs
Ordinary attn.
SAF 25%
Variance
Gate-Taylor
Adaptation: three pretraining seeds, task seed 1337
1,500
512 / 8
99.58±0.05
99.73±0.04
99.72±0.03
99.62±0.06
512 / 16
81.32±3.21
91.08±1.58
85.43±3.07
80.79±4.50
512 / 24
60.56±4.22
79.76±2.72
66.25±3.22
60.59±6.42
512 / 32
47.54±3.87
70.83±4.86
52.32±2.41
47.95±6.06
Appendix
Table 18: MQAR query accuracy (%). Adaptation varies one seed axis at a time (mean ± SE); immediate replacement and recovery use one task-trained ordinary-attention parent.
Update time (ms) ↓
Peak change (%)
Context
B
Rate
Ordinary attn.
Mean
Pruning
Mean
Pruning
4K
20
25%
2172.80
2057.35
1923.26
−2.15
−2.19
4K
20
50%
2173.12
1929.10
1678.24
−3.34
−4.22
8K
10
25%
2651.80
2573.68
2307.18
−2.24
−2.88
8K
10
50%
2652.83
2408.64
1933.85
−3.27
−4.95
16K
5
25%
3714.63
3621.11
3083.03
−1.76
−5.08
Appendix
Table 19: Native-context training-update costs for matched head sets. Memory changes are relative to ordinary attention at the same context and batch size.
Operation + selection
Ordinary attn. (ms)
Method (ms)
Speedup ( × )
Peak change (%)
Matched-token pretraining
Mean + variance
154.59
142.52
1.0847±0.0024
+3.75
Pruning + variance
155.26
133.43
1.1636±0.0012
−4.75
Pruning + gate-Taylor
154.22
132.63
1.1628±0.0009
−5.66
Pruning from initialisation
154.27
132.64
1.1631±0.0005
−4.75
Matched-time pretraining
Appendix
Table 20: Causal prefill at B=32 , T=4096 for the single-seed 25% controls. Speedup SE is over four timing pairs.
Task / checkpoint
B
Accuracy (%)
Δ accuracy (pp)
Time/update
Speedup
Peak memory change
SST-2, ordinary attn.
256
86.39±0.27
—
50.11±0.61 ms
1.000 ×
0.00%
SST-2, fixed 25%
256
86.54±0.36
+0.15±0.53
58.22±9.58 ms
0.901±0.122×
−3.65%
SST-2, fixed 50%
256
86.77±0.74
+0.38±1.00
56.70±9.43 ms
0.925±0.123×
−5.95%
BoolQ, ordinary attn.
64
71.40±0.22
—
67.36±0.51 ms
1.000 ×
0.00%
BoolQ, fixed 25%
64
70.62±0.55
−0.77±0.56
70.91±7.25 ms
0.967±0.083×
−4.63%
BoolQ, fixed 50%
64
70.84±0.49
−0.56±0.70
69.35±7.61 ms
0.992±0.093×
−7.19%
Appendix
Table 21: Finetuning quality, update time, and peak-memory change for inherited fixed-attention checkpoints.
Task
Schedule
Δ accuracy (pp)
Time/update
Controller cost
SST-2
One-time
−0.19±0.40
33.27 ms
0.70 s
Train-loss
−0.65±0.77
36.15 ms
45.12 s
Post-training
−1.03±0.24
30.75 ms
0.71 s
BoolQ
One-time
−1.27±1.29
40.85 ms
6.18 s
Train-loss
−0.38±0.15
66.34 ms
70.15 s
Post-training
−1.03±0.33
31.44 ms
6.17 s
Appendix
Table 22: Replacing 25% of heads under three RoPE finetuning schedules.
GPUs / B
Rate
Ordinary attn. (ms)
SAF (ms)
Speedup ( × )
Peak change (%)
Four GPUs, including communication: 524,288 tokens per update
4 / 2
25%
4258.67
3986.67
1.0682±0.0005
—
4 / 2
50%
4263.01
3653.78
1.1667±0.0003
—
Single GPU, no communication: 131,072 tokens per update
1 / 1
25%
4224.85
3958.26
1.067
−0.57
1 / 1
50%
4226.87
3632.25
1.163
−2.31
Appendix
Table 23: 1B post-replacement update cost at 8K. Four-GPU speedup errors are SE over four timing pairs; single-GPU results use medians.
Accuracy by seed ↑
Task
Model
1337
1338
1339
Change (pp)
Update (ms)
SST-2
Ordinary attention
91.63
90.94
91.74
—
173.58±0.74
SAF 25%
90.94
91.86
92.09
+0.19±0.47
166.78±2.75
Pruning 25%
91.06
91.86
92.20
+0.27±0.44
156.60±0.97
SAF 50%
91.28
91.51
90.14
−0.46±0.63
158.36±3.67
BoolQ
Ordinary attention
75.99
76.39
75.57
—
222.70±0.32
Appendix
Table 24: 1B finetuning by seed: accuracy (%), paired change from ordinary attention (pp), and observed mean update time. Errors are SE across finetuning seeds.
Attention variance
Forward KL
Task
Ordinary attn.
10%
20%
30%
10%
20%
30%
HellaSwag
52.28
52.02
51.21
48.92
49.10
45.05
43.81
PIQA
74.86
75.19
74.48
73.56
74.70
72.58
70.40
ARC-Easy
80.43
79.63
77.61
75.72
72.90
68.06
64.10
SST-2
88.65
88.65
87.96
85.44
87.61
63.07
68.81
BoolQ
84.46
83.33
80.49
73.12
80.03
71.38
71.38
Appendix
Table 25: Qwen3-4B zero-shot accuracy (%) with compact FineWeb-Edu mean patterns. The last row gives the mean change from the ordinary-attention model in percentage points.
B
T
10%
20%
30%
50%
75%
4
2048
1.011 (1.16)
1.028 (1.18)
1.068 (1.21)
1.135 (1.29)
1.225 (1.38)
4
4096
1.027 (1.16)
1.057 (1.19)
1.097 (1.23)
1.156 (1.28)
1.246 (1.38)
4
8192
1.042 (1.15)
1.072 (1.19)
1.115 (1.23)
1.184 (1.30)
1.250 (1.37)
8
2048
1.015 (1.15)
1.042 (1.17)
1.079 (1.21)
1.148 (1.28)
1.231 (1.37)
8
4096
1.029 (1.15)
1.061 (1.19)
1.080 (1.18)
1.156 (1.27)
1.252 (1.37)
8
8192
1.040 (1.15)
1.071 (1.18)
1.110 (1.22)
1.178 (1.30)
1.272 (1.40)
Appendix
Table 26: Qwen3-4B prefill speedup ( × ): end-to-end, with attention-operation speedup in parentheses.
Position
Ordinary attn. PPL
Dense PPL
Abs.+rel. PPL
Storage reduction
Absolute
26.0252
26.1184
26.1150
59.9 ×
RoPE
23.9571
24.1180
24.1165
60.6 ×
Appendix
Table 27: Representation quality at 25% replacement (one seed), prior storage, and kernel checks.
FP32
BF16 autocast
Tensor
Max. absolute
Relative L2
Max. absolute
Relative L2
Output
1.19×10−7
1.86×10−7
3.91×10−3
2.80×10−3
Input gradient dX
3.58×10−7
2.51×10−7
7.81×10−3
3.68×10−3
Value gradient dV
3.58×10−7
1.78×10−7
7.81×10−3
1.46×10−3
Query projection
8.34×10−7
2.82×10−7
1.56×10−2
4.53×10−3
Key projection
1.19×10−6
3.01×10−7
1.56×10−2
4.43×10−3
Appendix
Table 28: Numerical agreement with the eager reference for a layer with ordinary-attention and fixed-mean heads. Projection rows report weight gradients.
Input length
Batch size B
Replaced heads
Ordinary attn. (ms)
Replaced (ms)
Speedup
4K
64
25%
306.70
278.33
1.100 ×
4K
64
50%
303.94
247.47
1.229 ×
8K
64
25%
776.17
705.03
1.100 ×
8K
64
50%
760.20
622.94
1.221 ×
16K
64
25%
2075.04
1900.13
1.092 ×
16K
64
50%
2045.38
1703.87
1.200 ×
Appendix
Table 29: Selected causal-prefill latencies for the ordinary-attention and replaced 16K RoPE checkpoints at B=64 , without CUDA Graph replay. Input length varies while the checkpoints remain fixed.
Asset
Use
Upstream licence or terms
Datasets
FineWeb-Edu
Pretraining; calibration and evaluation
ODC-By 1.0 ; Common Crawl terms
SST-2 (GLUE)
Finetuning and zero-shot evaluation
Not specified in the original dataset card
BoolQ (SuperGLUE)
Finetuning and zero-shot evaluation
CC BY-SA 3.0
QuALITY
Long-input finetuning
CC BY 4.0 ; article-level licences
HellaSwag
Zero-shot evaluation
MIT
Appendix
Table 30: External datasets and pretrained models.
Prefilling computational costs pose a significant bottleneck for Large Language Models (LLMs) and Large Multimodal Models (LMMs) in long-context settings. While token pruning reduces sequence length, prior methods rely on heuristics that break compatibility with hardware-efficient kernels like FlashAttention. In this work, we observe that tokens evolve toward \textit{semantic fixing points}, making further processing redundant. To this end, we introduce Delta Attention Selective Halting (DASH), a training-free policy that monitors the layer-wise update dynamics of the self-attention mechanism to selectively halt stabilized tokens. Extensive evaluation confirms that DASH generalizes across language and vision benchmarks, delivering significant prefill speedups while preserving model accuracy and hardware efficiency. Code will be released at https://github.com/verach3n/DASH.git.
Yujie Chen, Tailai Chen, Yifeng Gao +4
Shanghai Jiao Tong University · University of California, San Diego · Carnegie Mellon University
Training causal transformers at extreme sequence lengths is bottlenecked by the quadratic time and memory of scaled dot-product attention (SDPA). In this work, we propose Lighthouse Attention, a training-only symmetrical selection-based hierarchical attention algorithm that wraps around ordinary SDPA and can be easily removed towards the end of the training. Our hierarchical selection is also gradient-free, which exempts us from dealing with a complicated and potentially inefficient backward pass kernel. Our contribution is three-fold: (i) A subquadratic hierarchical pre- and post-processing step that does adaptive compression and decompression of the sequence. (ii) A symmetrical compression strategy that pools queries, keys and values at the same time, while preserving left-to-right causality, which greatly improves parallelism. (iii) A two stage training approach which we pre-train for the majority of the time with Lighthouse Attention and recover a full attention model at the end with a short training. We run preliminary small scale LLM pre-training experiments that show the effectiveness of our method compared to full attention training with all other settings matched, where we achieve a faster total training time and lower final loss after the recovery phase. Full code is available at: https://github.com/ighoshsubho/lighthouse-attention
Long contexts have become standard in pretrained LLMs, yet they remain expensive to run: prefill compute grows quadratically with sequence length, and every decode step re-reads a key-value cache that grows linearly with it. Sparse attention cuts these costs by attending only to a relevant subset of past tokens, but selecting that subset is itself expensive. We present SpotAttention, a lightweight selector that attaches to a frozen pretrained transformer and learns by KL distillation to estimate its attention distribution. The selector picks the top-K keys each query attends to, and because its estimate is a calibrated distribution, a dual top-p rule reads the per-query, per-layer budget directly from it. Across Qwen3 (dense, 4B-32B) and Qwen3.5 (hybrid linear/full attention, 4B-9B), SpotAttention matches dense accuracy at contexts up to 128K tokens, eight times the training length. Decode at L=128K runs 3.9x faster than FlashAttention and 1.8x faster than Twilight, the strongest training-free baseline. Quantizing the selector's K-cache to INT4 or FP4 microscale shrinks it 3.5x at no accuracy cost.