Scaled dot-product attention forms its queries and keys as independent linear projections, so the two never interact before the dot product that scores them. We study coupled query-key dynamics, a pre-scoring transformation that evolves each token's query and key jointly through a shared invertible coupling before standard scoring. We realize it as an alternating affine map in the style of real non-volume-preserving flows: the coupling is the identity at initialization, adds a small fraction of parameters per head, and leaves the softmax and surrounding architecture unchanged. We place coupling on top of existing attention methods rather than replacing them, and ask whether that composition helps. On WikiText-103, adding coupling to Differential Attention improves on it at both 150M and 455M parameters. At 455M the gain is significant at sequence length 512 (p=0.003, six seeds), survives a Bonferroni correction and replicates on a held-out test split; it also holds across rotary-embedding training lengths 512, 1024, and 2048. The same additive direction appears when coupling is added to query-key normalization (significant at 150M) and Multi-Token Attention (directional). Matched controls attribute the gain to the joint pre-scoring coupling rather than to added capacity, and show that removing the invertibility guarantee preserves the 455M gain yet is far worse than the base method at 150M, so invertibility is what makes the coupling reliable across scales. On its own, coupling lowers perplexity at 60M and 150M (one-sided Welch tests, p<0.05) but the gain narrows with scale and is not significant at 455M. We relate the construction to the expressivity of coupling flows, use an associative-recall study to map where coupling helps and where it degrades sharp retrieval, and conclude that coupling is most useful in composition with a scoring-stage method rather than as a standalone change.
Figures & tables
Attention head
60M
150M
Standard
24.20±0.13
21.61±0.03
Affine coupling, one-way
24.04±0.02
21.53±0.05
Affine coupling, alternating
23.92±0.04
21.45±0.06
Table 1: The coupling family on WikiText-103, three seeds, mean ± std perplexity. Each step lowers perplexity; the alternating step is the strongly significant one at 60M ( p=0.006 ), while the intermediate steps are weaker or marginal (full ladder in Appendix B ).
60M
150M
455M
Standard
24.20±0.13
21.61±0.03
19.56±0.10
Alternating coupling
23.92±0.04
21.45±0.06
19.48±0.07
Relative
−1.18%
−0.74%
−0.36%
one-sided p
0.025
0.015
0.19
Table 2: Standalone coupling against standard attention across scale, three seeds, mean ± std. The advantage narrows and is not significant at 455M. Relative reductions are computed from per-seed perplexities, so they can differ slightly from the ratio of the rounded means shown here.
Method
Scale
Base
+ coupling
p
Differential Attention
150M
21.10±0.08
20.92±0.02
0.024
Differential Attention
455M
18.89±0.04
18.79±0.05
0.003
Multi-Token Attention
150M
20.96±0.05
20.85±0.12
0.128
Multi-Token Attention
455M
19.24±0.05
19.20±0.03
0.16
QK-Norm
150M
22.25±0.05
21.96±0.05
0.001
QK-Norm
455M
19.82±0.02
19.73±0.08
0.08
Table 3: Composition of the coupling with each base method at 150M and 455M, mean ± std perplexity with one-sided Welch p ; three seeds, except the 455M Differential Attention composition, which uses six. Each row pairs a method used alone (Base) against that method with the coupling added; bold marks a significant improvement ( p<0.05 ). The Differential Attention composition is significant at both scales and survives the Bonferroni correction; at 455M the Multi-Token Attention and query-key normalization compositions improve their base but are not significant, so the additive direction holds while the breadth narrows at scale. Full per-comparison statistics and the Bonferroni analysis are in Appendix B .
Pre-scoring variant
Params
60M perplexity
Standard
+0.00%
24.20±0.13
Independent nonlinear Q,K
+0.11%
24.13±0.03
Non-invertible cross MLP
+0.43%
24.04±0.07
Linear mix
+1.74%
24.34±0.05
Alternating coupling
+0.33%
23.92±0.04
Table 4: Capacity controls at 60M, three seeds, mean ± std perplexity, with parameter overhead over standard attention. Adding parameters without the invertible, joint coupling does not reproduce the gain, and the variants with the most parameters are not the best. The alternating coupling is significantly better than every control (one-sided Welch, p<0.05 ). Matched controls at 455M, including a non-invertible simultaneous variant of the coupling itself, are in the supplementary material.
Configuration
seed 1
seed 2
mean
Standard
45.34
45.25
45.30
Alternating coupling
44.66
44.89
44.78
Differential Attention
45.31
44.94
45.12
+ coupling
44.20
44.50
44.35
Table 5: FineWeb-Edu cross-corpus check at 60M. Per-seed and mean validation perplexity over two seeds; lower is better. Both seeds favor coupling in every row. With two seeds we report this as descriptive, with no significance test.
LAMBADA
HellaSwag
Configuration
150M
455M
150M
455M
Standard
0.259
0.254±0.006
0.313
0.303±0.004
Alternating coupling
0.269
0.252±0.003
0.317
0.297±0.006
Differential Attention
0.274
0.263±0.007
0.307
0.306±0.007
+ coupling
0.278
0.272±0.010
0.315
0.310±0.006
Table 6: Zero-shot transfer at 150M (single seed, seed 42) and 455M (mean ± std; six seeds per side for the Differential pair, three for standard and the standalone coupling); 2,000 examples per task. LAMBADA last-token accuracy and HellaSwag length-normalized accuracy; higher is better. A coupling variant is best in every column (bold). At 455M the composition is best on both tasks, and the standalone coupling falls below standard, consistent with its faded perplexity gain.
KV pairs
Seq. len
Standard
Coupling
Diff. Attn
+ Coupling
4
64
0.98±0.02
1.00±0.00
1.00
1.00
4
128
0.94±0.05
1.00±0.00
1.00
1.00
4
256
0.92±0.14
0.99±0.01
1.00
1.00
8
64
0.55±0.40
0.18±0.00
1.00
1.00
8
128
0.79±0.10
0.72±0.37
1.00
1.00
8
256
0.73±0.34
0.51±0.41
1.00
1.00
Table 7: Multi-query associative-recall accuracy, mean ± std over three seeds, on a small model ( d=256 , H=4 , L=6 ); higher is better. All variants reach 1.00 at two key-value pairs (row omitted). At four pairs the standalone coupling is higher and far more stable across seeds than standard attention; at eight pairs neither standalone head is reliable, and the coupling collapses at the densest cell. Composing either standalone head with Differential Attention restores recall to 1.00 on every cell (all ≥0.999 ).
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Config
dmodel
nheads
nlayers
dff
Params
Tiny
256
4
6
1024
∼ 6M †
Small
512
8
8
2048
∼ 60M
Medium
768
12
12
3072
∼ 153M
Large
1024
16
24
4096
∼ 455M
Appendix
Table 8: Model configurations used in experiments.
Variant
Hyperparameter
Value
Alternating coupling
Coupling rounds
1 (Eqs. ( 2 )–( 3 ))
Conditioning nets s,t
two-layer, SiLU, output width 2dk
Net placement
one (s,t) pair per update per layer, applied per head
Output-layer init
zero (exact identity at initialization)
Log-scale clamp
[−3,3]
GQA
KV head ratio
1/4 of query heads
Appendix
Table 9: Variant-specific hyperparameters. The coupling rows fully specify the alternating affine coupling of Eqs. ( 2 )–( 3 ).
Comparison
Scale
n
Baseline mean ± std
Treatment mean ± std
Δ
p
d
Note
1. Standalone coupling (vs. standard)
60M
3
24.20±0.13
23.92±0.04
−0.28
0.025∗
3.09
150M
3
21.61±0.03
21.45±0.06
−0.16
0.015∗
3.39
455M
3
19.56±0.10
19.48±0.07
−0.07
0.19
0.83
not significant
2. Composition: Differential Attention
150M
3
21.10±0.08
20.92±0.02
−0.17
0.024∗
3.12
455M
6
18.89±0.04
18.79±0.05
−0.10
0.003∗∗
2.09
455M
3
19.50±0.03
19.20±0.05
−0.30
0.0015∗∗
6.95
seq 1024, RoPE
Appendix
Table 10: Per-comparison statistics for all main claims. Δ= treatment PPL − baseline PPL (negative = improvement). Significance: ∗∗p<0.01 , ∗p<0.05 , †p<0.10 .
Comparison
n
Baseline
Treatment
p
Comp(Diff), seq 512
6
21.02±0.04
20.88±0.07
0.002
Comp(Diff), seq 1024, RoPE
3
20.13±0.08
19.81±0.06
0.003
Standalone, seq 512
3
21.79±0.05
21.76±0.02
0.21
Appendix
Table 11: WikiText-103 test-split perplexity of the validation-selected 455M checkpoints, mean ± std. Ordering and significance match the validation comparisons in Table 3 , including the non-significant standalone gap.
Scale
Pre-scoring variant on Diff. Attn.
Perplexity
Δ vs. base
p
455M
none (Differential Attention, 6 seeds)
18.89±0.04
455M
non-invertible cross-MLP (3 seeds)
18.89±0.11
+0.002
0.51
455M
simultaneous coupling (3 seeds)
18.78±0.07
−0.110
0.042
455M
invertible alternating coupling (6 seeds)
18.79±0.05
−0.097
0.003
150M
none (Differential Attention, 3 seeds)
21.10±0.08
150M
simultaneous coupling (3 seeds)
22.17±0.04
+1.078
0.0001 (worse)
Appendix
Table 12: Invertibility and capacity controls on top of Differential Attention, best validation perplexity, mean ± std. One-sided Welch p is against Differential Attention alone at the same scale (six seeds at 455M, three at 150M); for the two controls that are worse than the base, the reported p is for the reversed one-sided test. All variants are exactly their base method at initialization.
455M comparison
8K steps
15K steps
22K steps
29K steps
Composition vs. Differential Attention
−0.37%
−0.66%
−0.65%
−0.65%
Standalone coupling vs. standard
−0.31%
−0.73%
−0.79%
−0.34%
Appendix
Table 13: Best-so-far relative validation-perplexity gap (mean over three seeds; negative favors the coupling variant) as a function of training budget at 455M. The composition’s advantage is established early and remains stable to slightly widening through the end of training (late-half trend −0.02% per 1,000 steps), while the standalone gap closes late ( +0.02% per 1,000 steps), matching the faded standalone effect reported in the main paper.
Attention
Total Params
Variant-Specific
Overhead
Standard
60,343,296
0
n/a
Affine coupling, one-way
60,408,896
65,600
+ 0.11%
Affine coupling, alternating
60,541,952
198,656
+ 0.33%
GQA
57,197,568
0
− 5.21%
Diff Attention
60,343,360
64
+ 0.00%
Appendix
Table 14: Parameter comparison at 60M scale. The alternating affine coupling variant uses per-head coupling MLPs applied in an alternating pattern, adding approximately 0.3% parameters over standard attention.
Variant
Params
Rel. tok/s (512)
Rel. tok/s (2048)
Peak mem (MB)
Standard
+0.00%
1.00
1.00
4175
Differential Attention
+0.00%
0.86
0.68
5199
Alternating coupling
+0.33%
0.90
0.85
5071
+ Diff. Attention
+0.33%
0.79
0.61
6095
Appendix
Table 15: Measured wall-clock cost at 60M on one H100, batch 8, forward and backward. Throughput is tokens per second relative to standard attention (higher is better); peak memory is at sequence length 512. Because the per-head coupling MLPs are unfused, the realized cost exceeds the 0.33% parameter overhead. At length 2048 Differential Attention is itself the costlier component.
Variant
Params
Rel. tok/s (512)
Rel. tok/s (2048)
Mem 512 (GB)
Mem 2048 (GB)
Standard
+0.00%
1.00
1.00
16.7
26.4
Differential Attention
+0.00%
0.87
0.72
22.8
51.3
Alternating coupling
+0.13%
0.91
0.94
22.1
31.8
+ Diff. Attention
+0.13%
0.81
0.69
28.2
56.7
Appendix
Table 16: Measured wall-clock cost at 455M on one H100, forward and backward, batch 8 at sequence length 512 and batch 2 at length 2048. Throughput is tokens per second relative to standard attention at the same length; peak memory is reported at both lengths. The composition’s overhead over Differential Attention alone is 7.4% at length 512 and 4.1% at length 2048.
Variant
Post-softmax eff. rank (L16–L23 mean)
Standard
46.2
Alternating coupling
53.5
Differential Attention
42.5
+ coupling (composition)
45.3
Appendix
Table 17: Post-softmax attention-map effective rank, mean over L16–L23 of the 455M (24-layer) model (seed 42, WikiText-103). Pre-softmax logit rank is ≈14.4 for all variants (causal-mask dominated).
Attention
Learned Pos.
RoPE
Standard
24.20 ± 0.13
25.84 ± 0.07
Alternating coupling
23.92 ± 0.04
24.00 ± 0.03
Appendix
Table 18: RoPE ablation at 60M scale (alternating affine coupling). Mean ± std over 3 seeds.
The projection of queries and keys are central to the attention mechanism in Transformer architectures. While they are mathematically symmetric, they play different roles in attention mechanisms. The question of whether there is an effect from their functional distinction on their geometric development in training remains unanswered. We investigate the problem through the training of small GPT-like Transformers on character-level WikiText-103 for three different depths (4, 6, and 8 layers), three types of initialization for queries and keys, and four random seeds, resulting in 36 runs and 54 trajectories of average layers across seeds. We track the effective dimensionality of those layers using participation ratios and discover that effective dimension of queries expand while keys shrink, and that PRQ−PRK is positive in all trajectories studied. In connection to attention, the shrinking of keys leads to a narrower spectrum of QK⊤ and more peaked attention weights. In order to determine if this connection is causal or coincidental, we directly control the spectrum of keys during training across five seeds: restricting it to make it shrink sharpens the attention with high directional confidence, while keeping it constant to the level of initial dispersion makes attention softer. Additional token-level checkpoint analyses show that the monotonic paired-contrast trend is not universal across pretrained families, but survives as an early-training regime that later decays over a full pretraining run, and the link between interaction-rank geometry and attention entropy remains visible in several models.
We propose Keyless Attention, an attention mechanism that eliminates the key projection entirely, operating over queries and values only. This yields a Value-Only Cache that reduces KV cache memory and access overhead by exactly 50% over standard attention, while matching or exceeding standard attention's decode throughput. Beyond efficiency, we introduce Depth-m Attention Factorization: standard attention computes a depth-2 factorization of the attention bilinear form, while Keyless Attention realizes a depth-m instance of this family. At m=3, Keyless Attention matches the projection matrix count of standard attention via a value-space routing matrix that replaces the key projection and introduces a coupling between routing and retrieval. Experiments across five models and four architectures (GPT-2 280M, GPT-2 557M, Pythia 410M, Qwen2 1.5B, and Llama 3.2 1B) show that Keyless Attention matches or outperforms standard QKV attention on perplexity in 4 out of 5 models. On downstream zero-shot evaluation (GPT-2 557M), Keyless Attention outperforms on 4 out of 5 commonsense reasoning benchmarks, while achieving 50% KV cache reduction throughout.
Xin Gao
Department of Mathematics and Statistics York University Toronto, M3J 1P3
A standard self-attention layer consists of two interacting circuits: the query-key circuit that governs attention allocation, and the output-value circuit that maps attended representations to predictions. Collapsed and factorized parameterizations of the query-key and output-value circuits lead to qualitatively different attention patterns. In particular, some parameterizations give sharper attention to task-relevant tokens, at a similar training loss. We analyze how the parameterizations of these circuits shape the parameter trajectories in single-layer self-attention models trained for next-token prediction. Through gradient-flow analysis, we show that factorization induces implicit rescaling of the two circuits' learning rates. We derive closed-form dynamics showing that output-value and query-key parameters move along a line, with relative speeds determined by their learning rates. Faster query-key learning relative to output-value learning thus produces sharper attention, as the model compensates for slower output-value learning by increasing attention mass on relevant tokens. Experiments show that differences in the relative learning rates of the two circuits govern attention concentration. This improves attention interpretability proxies while maintaining comparable predictive performance.
Rahul Vashisht, Harish G. Ramaswamy
Department of CSE, IIT Madras, India · Department of DSAI, IIT Madras, India.