The projection of queries and keys are central to the attention mechanism in Transformer architectures. While they are mathematically symmetric, they play different roles in attention mechanisms. The question of whether there is an effect from their functional distinction on their geometric development in training remains unanswered. We investigate the problem through the training of small GPT-like Transformers on character-level WikiText-103 for three different depths (4, 6, and 8 layers), three types of initialization for queries and keys, and four random seeds, resulting in 36 runs and 54 trajectories of average layers across seeds. We track the effective dimensionality of those layers using participation ratios and discover that effective dimension of queries expand while keys shrink, and that PRQ−PRK is positive in all trajectories studied. In connection to attention, the shrinking of keys leads to a narrower spectrum of QK⊤ and more peaked attention weights. In order to determine if this connection is causal or coincidental, we directly control the spectrum of keys during training across five seeds: restricting it to make it shrink sharpens the attention with high directional confidence, while keeping it constant to the level of initial dispersion makes attention softer. Additional token-level checkpoint analyses show that the monotonic paired-contrast trend is not universal across pretrained families, but survives as an early-training regime that later decays over a full pretraining run, and the link between interaction-rank geometry and attention entropy remains visible in several models.
Figures & tables
Category
Metrics
Optimization
Training loss, validation loss
Attention Distribution
Attention entropy, top-1/top-5 attention mass, attention gap p(1)−p(2)
Query-Key Geometry
Cosine similarity between WQ and WK , column-wise alignment
Spectral Structure
Participation ratio (PR), effective rank, covariance eigenvalue statistics
Logit Geometry
Mean, variance, maximum attention logits
Norm and Scale
Activation norms, Frobenius norms, spectral norms
Table 1: Tracked geometric and optimization metrics.
Intervention
Manipulation
K Spectrum Clamp
Restores the singular-value spectrum of WK to its initialization spectrum every 100 optimizer steps from epoch 1 onward
Forced K Contraction
Shrinks the WK spectral tail after rank 32 by a factor of 0.05 every 100 optimizer steps, then renormalizes the projection
Freeze Q
Restores WQ to its initialization reference every optimizer step
Freeze K
Restores WK to its initialization reference every optimizer step
QK Decouple
Begins from equal Q/K initialization, then rotates WK once after epoch 3 begins
QK Align Regularized
Adds a weak Q/K alignment regularization penalty with λ=0.001
Table 2: Controlled interventions.
Figure 1: Initialization geometry and depth-wise attention roles. Top: WQ/WK alignment remains separated by initialization; bottom: depth-8 profiles reveal early, middle, and upper regimes.
Initialization
Runs
Final val loss
Final alignment
Depth-8 alignment
QK Identical
12
1.1907±0.0106
0.4530±0.0950
0.5048±0.0861
Disjoint Support Xavier
12
1.1969±0.0112
0.1991±0.0151
0.1911±0.0077
Independent Xavier
12
1.1973±0.0108
0.1895±0.0227
0.1799±0.0241
Table 3: Final loss and alignment summary.
Subset
n
PRQ slope
PRK slope
(PRQ−PRK) slope
All seed-averaged traces
54
+0.1406±0.0902
−0.0437±0.0565
+0.1843±0.0966
Uncoupled Q/K
36
+0.1366±0.0932
−0.0539±0.0547
+0.1904±0.0930
QK Identical
18
+0.1487±0.0859
−0.0235±0.0561
+0.1722±0.1051
Table 4: Participation-ratio slope summary.
Figure 2: Query participation outpaces key participation. Curves show seed-wise averaged PRQ−PRK trajectories; shaded bands show standard error across seeds.
Figure 3: Key specialization predicts interaction-spectrum compression. Compression, in turn, predicts entropy sharpening. Each point is one initialization-depth-layer cell with slopes averaged over the four seeds. Fitted lines and statistics use the 36 uncoupled cells; QK Identical cells are shown for reference.
Mode
Val loss
Entropy
Logit std
QK⊤ erank
Conf.
No intervention
1.2066
2.1893
4.4329
9.71
–
K spectrum clamp
1.2159
2.2888
3.5898
11.13
1.00
Forced K contraction
1.2270
2.0809
5.4071
6.03
1.00
Freeze Q
1.2196
2.5532
3.5915
9.35
0.93
Freeze K
1.2172
2.6002
3.2136
10.83
1.00
Q/K decouple
1.2115
1.7432
5.0723
9.92
0.93
Table 5: Seed-aggregated intervention outcomes across seeds 7, 37, 42, 105, and 210. Confidence is the fraction of seed-metric comparisons whose final-epoch delta has the same sign as the aggregate delta for entropy, logit standard deviation, and QK⊤ effective rank.
Figure 4: Five-seed intervention deltas against each seed’s matched no-intervention baseline. Bars show final-epoch mean deltas, error bars show 95% confidence intervals across seeds, and color indicates directional confidence score.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Family
Models
ΔPR slope >0
PRQ slope >0
PRK slope <0
Bridge co-move
Pythia
3
1/3
0/3
3/3
3/3
SmolLM2
2
0/2
0/2
0/2
0/2
GPT-2
1
0/1
1/1
0/1
1/1
All fitted series
6
1/6
1/6
3/6
4/6
Appendix
Table 6: Token-level checkpoint extension. Counts use only checkpoint series with at least three revisions. Bridge co-movement indicates positive correlation between PRK and QK⊤ effective rank together with positive correlation between QK⊤ effective rank and attention entropy.
Figure 5: Token-level checkpoint extension. The left panel shows fitted slopes of the normalized paired contrast (PRQ−PRK)/dhead across public checkpoint series. The right panel shows model-level correlations connecting key participation, QK⊤ effective rank, and attention entropy.