The projection of queries and keys are central to the attention mechanism in Transformer architectures. While they are mathematically symmetric, they play different roles in attention mechanisms. The question of whether there is an effect from their functional distinction on their geometric development in training remains unanswered. We investigate the problem through the training of small GPT-like Transformers on character-level WikiText-103 for three different depths (4, 6, and 8 layers), three types of initialization for queries and keys, and four random seeds, resulting in 36 runs and 54 trajectories of average layers across seeds. We track the effective dimensionality of those layers using participation ratios and discover that effective dimension of queries expand while keys shrink, and that PRQ−PRK is positive in all trajectories studied. In connection to attention, the shrinking of keys leads to a narrower spectrum of QK⊤ and more peaked attention weights. In order to determine if this connection is causal or coincidental, we directly control the spectrum of keys during training across five seeds: restricting it to make it shrink sharpens the attention with high directional confidence, while keeping it constant to the level of initial dispersion makes attention softer. Additional token-level checkpoint analyses show that the monotonic paired-contrast trend is not universal across pretrained families, but survives as an early-training regime that later decays over a full pretraining run, and the link between interaction-rank geometry and attention entropy remains visible in several models.
Figures & tables
Category
Metrics
Optimization
Training loss, validation loss
Attention Distribution
Attention entropy, top-1/top-5 attention mass, attention gap p(1)−p(2)
Query-Key Geometry
Cosine similarity between WQ and WK , column-wise alignment
Spectral Structure
Participation ratio (PR), effective rank, covariance eigenvalue statistics
Logit Geometry
Mean, variance, maximum attention logits
Norm and Scale
Activation norms, Frobenius norms, spectral norms
Table 1: Tracked geometric and optimization metrics.
Intervention
Manipulation
K Spectrum Clamp
Restores the singular-value spectrum of WK to its initialization spectrum every 100 optimizer steps from epoch 1 onward
Forced K Contraction
Shrinks the WK spectral tail after rank 32 by a factor of 0.05 every 100 optimizer steps, then renormalizes the projection
Freeze Q
Restores WQ to its initialization reference every optimizer step
Freeze K
Restores WK to its initialization reference every optimizer step
QK Decouple
Begins from equal Q/K initialization, then rotates WK once after epoch 3 begins
QK Align Regularized
Adds a weak Q/K alignment regularization penalty with λ=0.001
Table 2: Controlled interventions.
Figure 1: Initialization geometry and depth-wise attention roles. Top: WQ/WK alignment remains separated by initialization; bottom: depth-8 profiles reveal early, middle, and upper regimes.
Initialization
Runs
Final val loss
Final alignment
Depth-8 alignment
QK Identical
12
1.1907±0.0106
0.4530±0.0950
0.5048±0.0861
Disjoint Support Xavier
12
1.1969±0.0112
0.1991±0.0151
0.1911±0.0077
Independent Xavier
12
1.1973±0.0108
0.1895±0.0227
0.1799±0.0241
Table 3: Final loss and alignment summary.
Subset
n
PRQ slope
PRK slope
(PRQ−PRK) slope
All seed-averaged traces
54
+0.1406±0.0902
−0.0437±0.0565
+0.1843±0.0966
Uncoupled Q/K
36
+0.1366±0.0932
−0.0539±0.0547
+0.1904±0.0930
QK Identical
18
+0.1487±0.0859
−0.0235±0.0561
+0.1722±0.1051
Table 4: Participation-ratio slope summary.
Figure 2: Query participation outpaces key participation. Curves show seed-wise averaged PRQ−PRK trajectories; shaded bands show standard error across seeds.
Figure 3: Key specialization predicts interaction-spectrum compression. Compression, in turn, predicts entropy sharpening. Each point is one initialization-depth-layer cell with slopes averaged over the four seeds. Fitted lines and statistics use the 36 uncoupled cells; QK Identical cells are shown for reference.
Mode
Val loss
Entropy
Logit std
QK⊤ erank
Conf.
No intervention
1.2066
2.1893
4.4329
9.71
–
K spectrum clamp
1.2159
2.2888
3.5898
11.13
1.00
Forced K contraction
1.2270
2.0809
5.4071
6.03
1.00
Freeze Q
1.2196
2.5532
3.5915
9.35
0.93
Freeze K
1.2172
2.6002
3.2136
10.83
1.00
Q/K decouple
1.2115
1.7432
5.0723
9.92
0.93
Table 5: Seed-aggregated intervention outcomes across seeds 7, 37, 42, 105, and 210. Confidence is the fraction of seed-metric comparisons whose final-epoch delta has the same sign as the aggregate delta for entropy, logit standard deviation, and QK⊤ effective rank.
Figure 4: Five-seed intervention deltas against each seed’s matched no-intervention baseline. Bars show final-epoch mean deltas, error bars show 95% confidence intervals across seeds, and color indicates directional confidence score.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Family
Models
ΔPR slope >0
PRQ slope >0
PRK slope <0
Bridge co-move
Pythia
3
1/3
0/3
3/3
3/3
SmolLM2
2
0/2
0/2
0/2
0/2
GPT-2
1
0/1
1/1
0/1
1/1
All fitted series
6
1/6
1/6
3/6
4/6
Appendix
Table 6: Token-level checkpoint extension. Counts use only checkpoint series with at least three revisions. Bridge co-movement indicates positive correlation between PRK and QK⊤ effective rank together with positive correlation between QK⊤ effective rank and attention entropy.
Figure 5: Token-level checkpoint extension. The left panel shows fitted slopes of the normalized paired contrast (PRQ−PRK)/dhead across public checkpoint series. The right panel shows model-level correlations connecting key participation, QK⊤ effective rank, and attention entropy.
We propose Keyless Attention, an attention mechanism that eliminates the key projection entirely, operating over queries and values only. This yields a Value-Only Cache that reduces KV cache memory and access overhead by exactly 50% over standard attention, while matching or exceeding standard attention's decode throughput. Beyond efficiency, we introduce Depth-m Attention Factorization: standard attention computes a depth-2 factorization of the attention bilinear form, while Keyless Attention realizes a depth-m instance of this family. At m=3, Keyless Attention matches the projection matrix count of standard attention via a value-space routing matrix that replaces the key projection and introduces a coupling between routing and retrieval. Experiments across five models and four architectures (GPT-2 280M, GPT-2 557M, Pythia 410M, Qwen2 1.5B, and Llama 3.2 1B) show that Keyless Attention matches or outperforms standard QKV attention on perplexity in 4 out of 5 models. On downstream zero-shot evaluation (GPT-2 557M), Keyless Attention outperforms on 4 out of 5 commonsense reasoning benchmarks, while achieving 50% KV cache reduction throughout.
Xin Gao
Department of Mathematics and Statistics York University Toronto, M3J 1P3
FFT-based spectral preprocessing of learned query-key (Q/K) projections substantially improves transformer attention on character-level language modelling. On TinyShakespeare: a fixed random spectral filter achieves val=1.031 (Delta=+0.443); a single learned frequency at paragraph scale achieves val=0.608 (Delta=+0.867); and four learned frequencies spanning paragraph to word scale achieve val=0.309 (Delta=+1.166), a 79% reduction over standard dot-product attention. The single-frequency result is confirmed across three random seeds (mean val=0.236, std=0.019). The four frequencies converge to a near-geometric multi-scale ordering (49, 27, 10, 6 tokens/cycle) corresponding to paragraph, sub-paragraph, phrase, and word scales. The gain is specific to spectral preprocessing: random orthogonal and non-orthogonal projections of Q/K produce no measurable improvement, suggesting the benefit comes from global frequency-domain mixing rather than metric distortion. All results are verified by a shuffled-validation diagnostic against positional leakage. Causal filters (Gaussian, Mexican Hat, Morlet) do not improve over standard attention at character-level tokenisation: the bilateral FFT kernel is structurally non-causal, coupling every position to future tokens. This defines an architectural boundary between bilateral spectral attention (this paper) and genuinely causal spectral attention at word-scale tokenisation (companion paper MorletQK). This work is architecturally distinct from FNet (Lee-Thorp et al., 2021), which replaces attention with Fourier mixing of token embeddings. Here, spectral preprocessing applies only to Q/K projections while the full attention score structure is preserved.
Feed-forward networks hold two thirds of a transformer's non-embedding parameters, yet the architecture has not received a necessity test that controls parameters, compute, and depth at once. We pretrain attention-only decoder transformers (Simple Attention Networks, SANs) against standard transformers matched separately for parameter count, training FLOPs, and depth (2 to 48 layers), for up to 105B tokens at 6M to 87M parameters. Deleting feed-forward layers in place is costly: the standard transformer leads by 0.47 nats at matched depth and 0.26 nats at matched FLOPs. Reallocating the freed budget into attention depth closes the gap: at matched parameters the difference is 0.006 nats (0.27 percent of loss), reproducible to one part in ten thousand across seed pairs, shrinking across 5B, 30B, and 105B budgets, and holding near 0.02 nats across a 29x size range. Three measurements localize the remaining gap to parametric recall: attention-only models are better on context-grounded answers and worse where knowledge must come from weights. Weight spectra show why: routing matrices (Q/K) crystallize early, content matrices accumulate rank slowly, and removing feed-forward layers relocates this accumulation to the attention output projection. QK-normalization, not feed-forward layers or residual gating, keeps 48-layer attention-only stacks trainable. The deficit concentrates on low-context query prediction and localizes there entirely by the largest budget. A pre-registered test confirms the account: it predicts a 0.02 to 0.05 nat gap on knowledge-dense web text; a matched pair trained on fineweb-edu measures 0.040. Within the tested regime, attention does the rest.