The query and key projections
\WQ,\WK in attention are almost always trained by Euclidean optimizers with no geometric constraint. We constrain them to the Stiefel manifold and optimize with a Riemannian Adam carrying one scalar second moment per frame---the form of \citet{becigneul2019}, here extended to the compact, non-Hadamard
\St(d,r) with a tangent projector, step-norm cap, and polar retraction. Four propositions prove steepest descent in the embedded metric, gradient-scale independence, well-conditioning, and exact
O(d)-equivariance. A fifth records that weight decay has \emph{identically zero} Riemannian gradient on
\St(d,r) (
W=WIr lies in the normal space), so decay cannot act on the constrained frames. On a CIFAR-10 patch benchmark at
n=10k this rule gains
+6.79,pp over AdamW across 12 paired starts (
t=38.33,
12/12); earlier fixed-step Riemannian SGD gains
+1.97,pp, of which
+1.69,pp comes from frozen orthonormal initialization alone. The corrected Adam's lead grows with data:
+1.9,pp at
n=1k to
+6.7,pp at
n=50k. A 12-seed ablation credits all gain to the scale-free step (
+4.63,pp,
12/12), nothing to the projector or equivariance; a targeted
ε-sweep causally confirms the mechanism (
−2.6,pp at
ε=0.1,
p<0.001). Two five-seed grokking studies confirm the constrained arm does not grok better than the baseline (
p=0.019, A2 wins): the weight-decay exemption has no grokking consequence. A single-seed pilot exploiting this localization achieves the first stable grokking under slingshot conditions---Stiefel + targeted circuit regularization keeps routing-frame isometry error
106× lower than the unconstrained ablation through every collapse.