As data propagates through a Transformer, the norm of its hidden states grows by orders of magnitude with depth, a phenomenon framed as 'curse of depth' and nearly universally treated as a pathology to be suppressed. We take the opposite view. Across 16 pre-trained LLMs from 9 families, spanning dense, mixture-of-experts and hybrid architectures and Pre-, Peri- and Post-Norm designs, we find that this growth reflects an emergent depth-positional encoding, carried by the only learned per-layer gain on the residual stream, the normalization weight γ: with depth, γ grows in magnitude and rotates in direction, jointly encoding the layer index. We make this depth-conditioned encoding explicit with LayerRoPE, an implicit analog of RoPE along the depth axis, which replaces all layerwise γ vectors with a single shared vector and depth-conditioned scalars, at a net reduction in parameters and <0.02% change in FLOPs. Across a model ladder scaled up to 100B+ tokens, LayerRoPE consistently outperforms Pre-, Post- and Peri-Norm and Layer-Norm Scaling, reaching Pre-Norm's 1.3B loss with 3.4× less compute; LayerRoPE is the only approach that shows strong convergence and improves near monotonically as depth scales to 512 layers. It improves learning-rate sensitivity by 3-10×, and transfers naively to and consistently improves looped latent models and Vision Transformers. Inspecting its learned schedule inverts the prevailing premise: LayerRoPE does not shrink the residual stream but widens it, damping what each block reads while amplifying what it writes. Depth stability, our results suggest, calls not for suppressing the residual stream, but for depth-conditioned regulation of the computational blocks it feeds.
Figures & tables
Figure 1: LayerRoPE offers a contrary premise: Maximizing residual variance with depth-conditioned regulation of computational blocks , allows better convergence, deeper networks, wider learning basins, efficiently ( <0.02% FLOP increase), and across architectures and modalities.
Figure 2: Norm weights encode depth in LLAMA-like LLMs (LLaMA-2 7B shown here). Magnitude (left) and angular position (right) jointly form a depth-positional code along the residual stream.
Figure 3: Per-dimension RMSNorm gain γ across layers, for the Attention-input (left panel) and MLP-input (right panel) norms of a particular head (Head 3), in LLaMA-2 7B. Almost all dimensions share a single, near-monotone depth-conditioned trajectory.
Accuracy (%, ↑ )
Method
ARC-e
HellaS.
PIQA
BoolQ
OBQA
Wino.
LMB.
MMLU
Avg ( ↑ )
LMB PPL ( ↓ )
134M [2pt]10.7 B
Pre-Norm
35.0
31.1
63.9
56.2
28.8
48.2
22.6
23.0
40.8
110.5
Post-Norm
36.2
30.1
62.1
58.8
27.2
51.1
21.1
23.0
41.0
125.1
Peri-Norm
36.6
31.5
63.0
60.6
27.8
51.3
24.4
22.9
42.2
0 83.1
LNS
37.1
32.0
63.2
57.3
27.4
48.8
23.1
22.9
41.3
0 93.8
LayerRoPE
37.8
32.6
64.1
59.6
28.6
51.2
25.9
23.0
42.8
0 68.5
Table 1: Downstream Evaluations on scaling ladder. Zero-shot accuracy of all methods at their tuned learning rate. Within each scale, best and second-best are marked per column. † indicates outlier Lambada perplexity ( >104 ). (-) indicate Post-Norm failed to converge well at 1.3B scale.
Figure 4: Model Ladder Learning-rate Sweep. Each method is swept on the shared ×2 grid until its optimum is interior; outlined markers indicate the tuned optima, whose rates are listed in Table 6 .
Figure 5: Depth scaling. Final validation loss against the number of layers. Points are means over two seeds and bars span both seeds (up to 384 layers). The upper strip shows diverged Post-Norm runs.
Figure 6: Learning Rate Basins. Final validation loss against peak learning rate for Pre-Norm (left) and Peri-Norm (right), each with and without LayerRoPE, at 58M, 368M and 1.3B scales. Runs with losses too high for the panel or diverged, are omitted.
Figure 8
Acc. (%) ↑
Loss ↓
Model
Top-1
Top-5
Val.
ViT-T
72.21
91.26
1.211
+ LayerRoPE
73.08
91.62
1.171
ViT-S
79.80
94.89
0.897
+ LayerRoPE
80.38
95.26
0.865
Table 3: ViTs on ImageNet1k. Each metric is measured as the best over training.
Figure 8: LayerRoPE widens the residual stream. Activation variance of the residual stream with model depth, across 368M and 1.3B models, measured at the output of each layer, on the C4 validation set.
Table 4: Ablating LayerRoPE at 134M. Validation loss and perplexity, best in bold ; Baseline is the backbone without LayerRoPE. (a) Removing either schedule degrades LayerRoPE on Peri-Norm, though each alone still improves on the baseline. (b) Separate gains γ per layer add parameters without improving on LayerRoPE’s single shared γ , while one rotation angle for all feature pairs underperforms RoPE’s spectrum of frequencies.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Pre-Norm
FLOPs overhead (%)
LayerRoPE param reduction
Model
d
L
GFLOPs/token
vs. Pre-Norm
vs. Peri-Norm
vs. Pre-Norm
vs. Peri-Norm
60M
512
8
0.26
0.0128
≈0
↓ 6.128K
↓ 14.320K
130M
768
12
0.67
0.0110
≈0
↓ 15.344K
↓ 33.776K
250M
896
20
1.36
0.0106
≈0
↓ 32.240K
↓ 68.080K
500M
1280
24
3.12
0.0079
≈0
↓ 56.304K
↓ 117.744K
1B
2048
24
7.72
0.0051
≈0
↓ 90.096K
↓ 188.400K
Appendix
Table 5: Estimated Analytic training FLOPs of the Pre-Norm baseline, and the relative overhead (%) and parameter reduction ( green ↓ ) of LayerRoPE over Pre-Norm and Peri-Norm. Over Peri-Norm, LayerRoPE adds only the negligible cost ( <10−4% ) of constructing the depth-conditioned γℓ , shown as ≈0 .
Figure 9: (Pre-Norm, MLP Blocks) Depth-conditioning of Norm Weights. For each pre-trained model ( GPT-OSS 20B, Mistral 7B, Pythia 2.8B and LLaMA-3 8B ), the distribution of MLP Block RMSNorm γ values and MDS Projection of the same per-layer γ vectors is shown. Note that in all cases, we see strong depth-conditioning in both layer-wise angular and magnitude properties of MLP norm weights.
Figure 10: (Pre-Norm, MLP Blocks. (continued) ) Depth-conditioning of Norm Weights. For each pre-trained model ( LLaMA-1 7B, Qwen3 8B and Phi-2 2.7B ), we continue to see strong depth-conditioning in both layer-wise angular and magnitude properties of the MLP norm weights. We note Qwen2.5, Qwen3.5 and Gemma (highlighted in the boxed row), unlike all other Pre-Norm models, do not show a strong depth conditioning effect in their MLP Norm weights.
Figure 11: Depth encoding in the Attention block norms of Pre-Norm LLMs. For each model, top : MDS projection of the per-layer Attention block γ vectors (angle); bottom : per-layer distribution of the same γ values (magnitude). Compare with the MLP block norms of the same models in Figure 9 (LLaMA-2’s in Figure 2 ).
Figure 12: (Peri-Norm, MLP Blocks) Depth-conditioning of Norm Weights. For each pre-trained model ( Gemma 2 9B and GLM-4 9B ), the input and output norms of the MLP block are shown: (a, b) Gemma 2 9B and (c, d) GLM-4 9B. In each panel, the MDS projection of the per-layer γ vectors (angle) sits above the distribution of γ values (magnitude); Gemma 2’s gains are shown as the applied 1+γ . We omit the Attention block norms, where as expected, the depth-conditioning is weak
Figure 13: (Post-Norm, MLP and Attention Blocks) Depth-conditioning of Norm Weights. For each pre-trained model ( OLMo 2 7B and OLMo 2 13B ), the MLP and Attention block norms are shown. OLMo 2 normalizes only the output of each block, so each model has an MLP and an Attention output norm. Within each panel, the MDS projection (angle) sits above the distribution of γ values (magnitude).
Figure 14: Per-dimension RMSNorm gain γ across layers (filtered by specific Heads), for the Attention-input (left panel of each row) and MLP-input (right panel) norms of LLaMA-2 7B.
Peak learning rate ( ×10−3 )
Model
Params
Tokens
C (FLOPs)
Pre
Post
Peri
LNS
LayerRoPE
60M
58.1M
4.6B
1.2×1018
16
2
32
16
64
130M
134.1M
10.7B
7.2×1018
8
2
8
16
32
250M
250.0M
20.0B
2.7×1019
8
1
16
16
64
500M
553.8M
44.3B
1.4×1020
4
0.0625
16
8
32
1B
1.34B
107.1B
8.3×1020
2
0.031 †
8
8
32
Appendix
Table 6: Scaling ladder and tuned peak learning rates. Tokens and C are the total training tokens and compute of each run (Appendix D.1.3 ). Rates for 60M-500M are measured optima; rates for 1B are single runs at each method’s extrapolated optimum (Appendix D.1.2 ). † Failed to converge.
Final validation loss
Paired gap
Params
Pre
Post
Peri
Layer-Norm Scaling
LayerRoPE
LayerRoPE − Pre
58M
3.172 ± 0.001
3.201 ± 0.001
3.158 ± 0.014
3.147 ± 0.002
3.129 ± 0.016
−0.043± 0.016
134M
2.915 ± 0.014
2.941 ± 0.003
2.884 ± 0.010
2.859 ± 0.002
2.831 ± 0.001
−0.085± 0.014
Appendix
Table 7: Seed-to-seed variability on the scaling ladder. Final validation loss of every method at its tuned learning rate (Table 6 ), as mean ± standard deviation over three seeds. A seed sets both the initialization and the data order and is shared by all methods, so the last column is a paired difference.
Figure 15: Extrapolating the tuned learning rate to 1B. Filled marks are each method’s tuned optima at the four swept scales (Figure 4 ). Lines are the per-method least-squares fits logLR⋆=am+bmlogP , solid over the fitted range and dashed where extrapolated. At the 1B tier ( P=1.34 B), each small dot is a fit’s prediction and each hollow marker is that prediction rounded to the nearest rate on the ×2 grid (horizontal lines); the hollow markers are the rates used for the 1B runs (Table 6 ).
C4 validation
Accuracy (%, ↑ )
Method
Loss ( ↓ )
PPL ( ↓ )
ARC-e
HellaS.
PIQA
BoolQ
OBQA
Wino.
LMB.
MMLU
Avg ( ↑ )
LMB PPL ( ↓ )
Pre-Norm
2.545
12.74
42.7
40.5
67.8
60.3
31.0
51.5
29.4
23.0
46.2
37.6
Post-Norm
7.172
1,302
26.2
26.5
49.0
37.8
24.4
48.1
0 0.0
22.9
30.3
- †
Peri-Norm
2.514
12.35
42.3
41.7
67.9
58.3
30.8
50.4
32.2
23.0
46.2
30.8
Layer-Norm Scaling
2.464
11.75
45.4
45.6
70.9
49.9
30.8
52.6
34.2
23.0
47.0
28.1
LayerRoPE
2.377
10.77
44.0
49.6
70.5
56.5
31.0
56.0
37.6
22.9
49.3
19.0
Appendix
Table 8: Scaling to 6.7B parameters. Final C4 validation loss and perplexity, and zero-shot accuracy on the tasks of Table 1 , after 11.7B training tokens at a shared learning rate. Best and second-best are marked per column. † Lambada perplexity above 104 .
Figure 16: Learning-rate sweeps at every Depth. Final validation loss against peak learning rate for the six methods of Figure 5 at each depth. Outlined markers indicate each method’s best rate; loss axes are set per panel.
Figure 17: Depth scaling at the tuned learning rate. Final validation loss against the number of layers, with every method at its best rate from Figure 16 (seed 1). At 512 layers, Post-Norm diverged at every rate (upper strip).
η⋆
L⋆
Largest stable rate
LR sensitivity S
Backbone
Params
−
+
−
+
−
+
−
+
Pre-Norm
58M
16
64
3.172
3.118
16
128 ‡
0.458
0.119
368M
4
20
2.759
2.660
8
50
1.069
0.360
1.3B
0.5
1
2.648
2.505
0.5
8 ‡
1.162
0.112
Peri-Norm
58M
32
64
3.152
3.120
32
128 ‡
0.440
0.087
368M
8
8
2.702
2.671
8
100 ‡
1.310
0.263
Appendix
Table 9: Learning-rate robustness in the dense sweeps (Figure 6 ), without ( − ) and with ( + ) LayerRoPE: the tuned rate η⋆ and loss L⋆ , the largest rate that trained stably, and LR sensitivity S ( Wortsman et al., 2024 ) , the mean of min(L(η),L0)−minηL(η) over the rates swept for both arms, with L0=ln32000 . Rates are in units of 10−3 . ‡ Largest rate swept.
Figure 18: Every loop depth at 770M. Validation loss throughout training for Parcae (solid) and Parcae + LayerRoPE ( dashed ), evaluated at T∈{1,2,4,8} loop iterations, from lighter to darker. The loss axis omits the first ∼1 B tokens.
Figure 19: Recurrent state during training. Mean norm of the recurrent state leaving the recurrent block (top, log scale) and mean pairwise cosine similarity between its centred token states (bottom), for Parcae and Parcae + LayerRoPE. Lines are rolling medians over logged steps, with raw values shown faintly.
Accuracy (%, ↑ )
Model
Looped Iters. (T)
Val. PPL ( ↓ )
Lambada PPL ( ↓ )
ARC-e
HellaS.-0
HellaS.-10
PIQA
LMB.
SQuAD
CoQA
QA-WD
Avg. ( ↑ )
140M [2pt]11.2 B
GPT
-
20.28
105.61
51.0
31.6
31.7
62.2
24.3
5.0
9.8
23.3
29.9
Parcae
8
18.66
81.17
53.6
33.7
33.7
63.1
26.4
7.5
12.8
33.1
33.0
Parcae + LayerRoPE
8
18.36
67.32
54.0
34.3
34.5
63.9
27.1
14.0
15.6
35.7
34.9
370M [2pt]29.6 B
GPT
-
15.01
37.16
59.5
40.3
40.1
67.0
32.9
20.2
20.1
44.4
40.6
Parcae
8
14.22
32.46
61.4
43.6
43.5
68.5
35.0
28.0
20.6
46.8
43.4
Appendix
Table 10: Per-task downstream results of the looped latent models. Validation perplexity on FineWeb-Edu, Lambada perplexity and accuracy (%) on each downstream task, averaged over three few-shot sampling seeds, with looped models at T=8 ; training tokens under each size, best per size in bold .
Figure 20: Learned depth multipliers at 1.3B. Learned magnitude multipliers er(ℓ) of LayerRoPE against the fixed 1/ℓ+1 of Layer-Norm Scaling.
Figure 21: Learned rotation angles at 1.3B. Rotation θ(ℓ) of the highest-frequency channel pair at each LayerRoPE site. Labels give the learned slopes βrot .
Figure 22: Effective depth gain at 1.3B. Magnitude ratio ∥γℓeff∥/∥γ0eff∥ of each layer’s effective normalization weights to the first layer’s, with every depth-dependent factor folded into γℓeff (legend). For Peri-Norm, the right column shows its output normalization; Pre-Norm and Layer-Norm Scaling have no residual-write weights. LayerRoPE’s weights γ are shared across layers, so its ratio is (ℓ+1)βmag ; labels give the learned slopes βmag .