As data propagates through a Transformer, the norm of its hidden states grows by orders of magnitude with depth, a phenomenon framed as 'curse of depth' and nearly universally treated as a pathology to be suppressed. We take the opposite view. Across 16 pre-trained LLMs from 9 families, spanning dense, mixture-of-experts and hybrid architectures and Pre-, Peri- and Post-Norm designs, we find that this growth reflects an emergent depth-positional encoding, carried by the only learned per-layer gain on the residual stream, the normalization weight γ: with depth, γ grows in magnitude and rotates in direction, jointly encoding the layer index. We make this depth-conditioned encoding explicit with LayerRoPE, an implicit analog of RoPE along the depth axis, which replaces all layerwise γ vectors with a single shared vector and depth-conditioned scalars, at a net reduction in parameters and <0.02% change in FLOPs. Across a model ladder scaled up to 100B+ tokens, LayerRoPE consistently outperforms Pre-, Post- and Peri-Norm and Layer-Norm Scaling, reaching Pre-Norm's 1.3B loss with 3.4× less compute; LayerRoPE is the only approach that shows strong convergence and improves near monotonically as depth scales to 512 layers. It improves learning-rate sensitivity by 3-10×, and transfers naively to and consistently improves looped latent models and Vision Transformers. Inspecting its learned schedule inverts the prevailing premise: LayerRoPE does not shrink the residual stream but widens it, damping what each block reads while amplifying what it writes. Depth stability, our results suggest, calls not for suppressing the residual stream, but for depth-conditioned regulation of the computational blocks it feeds.
Figures & tables
Figure 1: LayerRoPE offers a contrary premise: Maximizing residual variance with depth-conditioned regulation of computational blocks , allows better convergence, deeper networks, wider learning basins, efficiently ( <0.02% FLOP increase), and across architectures and modalities.
Figure 2: Norm weights encode depth in LLAMA-like LLMs (LLaMA-2 7B shown here). Magnitude (left) and angular position (right) jointly form a depth-positional code along the residual stream.
Figure 3: Per-dimension RMSNorm gain γ across layers, for the Attention-input (left panel) and MLP-input (right panel) norms of a particular head (Head 3), in LLaMA-2 7B. Almost all dimensions share a single, near-monotone depth-conditioned trajectory.
Accuracy (%, ↑ )
Method
ARC-e
HellaS.
PIQA
BoolQ
OBQA
Wino.
LMB.
MMLU
Avg ( ↑ )
LMB PPL ( ↓ )
134M [2pt]10.7 B
Pre-Norm
35.0
31.1
63.9
56.2
28.8
48.2
22.6
23.0
40.8
110.5
Post-Norm
36.2
30.1
62.1
58.8
27.2
51.1
21.1
23.0
41.0
125.1
Peri-Norm
36.6
31.5
63.0
60.6
27.8
51.3
24.4
22.9
42.2
0 83.1
LNS
37.1
32.0
63.2
57.3
27.4
48.8
23.1
22.9
41.3
0 93.8
LayerRoPE
37.8
32.6
64.1
59.6
28.6
51.2
25.9
23.0
42.8
0 68.5
Table 1: Downstream Evaluations on scaling ladder. Zero-shot accuracy of all methods at their tuned learning rate. Within each scale, best and second-best are marked per column. † indicates outlier Lambada perplexity ( >104 ). (-) indicate Post-Norm failed to converge well at 1.3B scale.
Figure 4: Model Ladder Learning-rate Sweep. Each method is swept on the shared ×2 grid until its optimum is interior; outlined markers indicate the tuned optima, whose rates are listed in Table 6 .
Figure 5: Depth scaling. Final validation loss against the number of layers. Points are means over two seeds and bars span both seeds (up to 384 layers). The upper strip shows diverged Post-Norm runs.
Figure 6: Learning Rate Basins. Final validation loss against peak learning rate for Pre-Norm (left) and Peri-Norm (right), each with and without LayerRoPE, at 58M, 368M and 1.3B scales. Runs with losses too high for the panel or diverged, are omitted.
Figure 8
Acc. (%) ↑
Loss ↓
Model
Top-1
Top-5
Val.
ViT-T
72.21
91.26
1.211
+ LayerRoPE
73.08
91.62
1.171
ViT-S
79.80
94.89
0.897
+ LayerRoPE
80.38
95.26
0.865
Table 3: ViTs on ImageNet1k. Each metric is measured as the best over training.
Figure 8: LayerRoPE widens the residual stream. Activation variance of the residual stream with model depth, across 368M and 1.3B models, measured at the output of each layer, on the C4 validation set.
Table 4: Ablating LayerRoPE at 134M. Validation loss and perplexity, best in bold ; Baseline is the backbone without LayerRoPE. (a) Removing either schedule degrades LayerRoPE on Peri-Norm, though each alone still improves on the baseline. (b) Separate gains γ per layer add parameters without improving on LayerRoPE’s single shared γ , while one rotation angle for all feature pairs underperforms RoPE’s spectrum of frequencies.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Pre-Norm
FLOPs overhead (%)
LayerRoPE param reduction
Model
d
L
GFLOPs/token
vs. Pre-Norm
vs. Peri-Norm
vs. Pre-Norm
vs. Peri-Norm
60M
512
8
0.26
0.0128
≈0
↓ 6.128K
↓ 14.320K
130M
768
12
0.67
0.0110
≈0
↓ 15.344K
↓ 33.776K
250M
896
20
1.36
0.0106
≈0
↓ 32.240K
↓ 68.080K
500M
1280
24
3.12
0.0079
≈0
↓ 56.304K
↓ 117.744K
1B
2048
24
7.72
0.0051
≈0
↓ 90.096K
↓ 188.400K
Appendix
Table 5: Estimated Analytic training FLOPs of the Pre-Norm baseline, and the relative overhead (%) and parameter reduction ( green ↓ ) of LayerRoPE over Pre-Norm and Peri-Norm. Over Peri-Norm, LayerRoPE adds only the negligible cost ( <10−4% ) of constructing the depth-conditioned γℓ , shown as ≈0 .
Figure 9: (Pre-Norm, MLP Blocks) Depth-conditioning of Norm Weights. For each pre-trained model ( GPT-OSS 20B, Mistral 7B, Pythia 2.8B and LLaMA-3 8B ), the distribution of MLP Block RMSNorm γ values and MDS Projection of the same per-layer γ vectors is shown. Note that in all cases, we see strong depth-conditioning in both layer-wise angular and magnitude properties of MLP norm weights.
Figure 10: (Pre-Norm, MLP Blocks. (continued) ) Depth-conditioning of Norm Weights. For each pre-trained model ( LLaMA-1 7B, Qwen3 8B and Phi-2 2.7B ), we continue to see strong depth-conditioning in both layer-wise angular and magnitude properties of the MLP norm weights. We note Qwen2.5, Qwen3.5 and Gemma (highlighted in the boxed row), unlike all other Pre-Norm models, do not show a strong depth conditioning effect in their MLP Norm weights.
Figure 11: Depth encoding in the Attention block norms of Pre-Norm LLMs. For each model, top : MDS projection of the per-layer Attention block γ vectors (angle); bottom : per-layer distribution of the same γ values (magnitude). Compare with the MLP block norms of the same models in Figure 9 (LLaMA-2’s in Figure 2 ).
Figure 12: (Peri-Norm, MLP Blocks) Depth-conditioning of Norm Weights. For each pre-trained model ( Gemma 2 9B and GLM-4 9B ), the input and output norms of the MLP block are shown: (a, b) Gemma 2 9B and (c, d) GLM-4 9B. In each panel, the MDS projection of the per-layer γ vectors (angle) sits above the distribution of γ values (magnitude); Gemma 2’s gains are shown as the applied 1+γ . We omit the Attention block norms, where as expected, the depth-conditioning is weak
Figure 13: (Post-Norm, MLP and Attention Blocks) Depth-conditioning of Norm Weights. For each pre-trained model ( OLMo 2 7B and OLMo 2 13B ), the MLP and Attention block norms are shown. OLMo 2 normalizes only the output of each block, so each model has an MLP and an Attention output norm. Within each panel, the MDS projection (angle) sits above the distribution of γ values (magnitude).
Figure 14: Per-dimension RMSNorm gain γ across layers (filtered by specific Heads), for the Attention-input (left panel of each row) and MLP-input (right panel) norms of LLaMA-2 7B.
Peak learning rate ( ×10−3 )
Model
Params
Tokens
C (FLOPs)
Pre
Post
Peri
LNS
LayerRoPE
60M
58.1M
4.6B
1.2×1018
16
2
32
16
64
130M
134.1M
10.7B
7.2×1018
8
2
8
16
32
250M
250.0M
20.0B
2.7×1019
8
1
16
16
64
500M
553.8M
44.3B
1.4×1020
4
0.0625
16
8
32
1B
1.34B
107.1B
8.3×1020
2
0.031 †
8
8
32
Appendix
Table 6: Scaling ladder and tuned peak learning rates. Tokens and C are the total training tokens and compute of each run (Appendix D.1.3 ). Rates for 60M-500M are measured optima; rates for 1B are single runs at each method’s extrapolated optimum (Appendix D.1.2 ). † Failed to converge.
Final validation loss
Paired gap
Params
Pre
Post
Peri
Layer-Norm Scaling
LayerRoPE
LayerRoPE − Pre
58M
3.172 ± 0.001
3.201 ± 0.001
3.158 ± 0.014
3.147 ± 0.002
3.129 ± 0.016
−0.043± 0.016
134M
2.915 ± 0.014
2.941 ± 0.003
2.884 ± 0.010
2.859 ± 0.002
2.831 ± 0.001
−0.085± 0.014
Appendix
Table 7: Seed-to-seed variability on the scaling ladder. Final validation loss of every method at its tuned learning rate (Table 6 ), as mean ± standard deviation over three seeds. A seed sets both the initialization and the data order and is shared by all methods, so the last column is a paired difference.
Figure 15: Extrapolating the tuned learning rate to 1B. Filled marks are each method’s tuned optima at the four swept scales (Figure 4 ). Lines are the per-method least-squares fits logLR⋆=am+bmlogP , solid over the fitted range and dashed where extrapolated. At the 1B tier ( P=1.34 B), each small dot is a fit’s prediction and each hollow marker is that prediction rounded to the nearest rate on the ×2 grid (horizontal lines); the hollow markers are the rates used for the 1B runs (Table 6 ).
C4 validation
Accuracy (%, ↑ )
Method
Loss ( ↓ )
PPL ( ↓ )
ARC-e
HellaS.
PIQA
BoolQ
OBQA
Wino.
LMB.
MMLU
Avg ( ↑ )
LMB PPL ( ↓ )
Pre-Norm
2.545
12.74
42.7
40.5
67.8
60.3
31.0
51.5
29.4
23.0
46.2
37.6
Post-Norm
7.172
1,302
26.2
26.5
49.0
37.8
24.4
48.1
0 0.0
22.9
30.3
- †
Peri-Norm
2.514
12.35
42.3
41.7
67.9
58.3
30.8
50.4
32.2
23.0
46.2
30.8
Layer-Norm Scaling
2.464
11.75
45.4
45.6
70.9
49.9
30.8
52.6
34.2
23.0
47.0
28.1
LayerRoPE
2.377
10.77
44.0
49.6
70.5
56.5
31.0
56.0
37.6
22.9
49.3
19.0
Appendix
Table 8: Scaling to 6.7B parameters. Final C4 validation loss and perplexity, and zero-shot accuracy on the tasks of Table 1 , after 11.7B training tokens at a shared learning rate. Best and second-best are marked per column. † Lambada perplexity above 104 .
Figure 16: Learning-rate sweeps at every Depth. Final validation loss against peak learning rate for the six methods of Figure 5 at each depth. Outlined markers indicate each method’s best rate; loss axes are set per panel.
Figure 17: Depth scaling at the tuned learning rate. Final validation loss against the number of layers, with every method at its best rate from Figure 16 (seed 1). At 512 layers, Post-Norm diverged at every rate (upper strip).
η⋆
L⋆
Largest stable rate
LR sensitivity S
Backbone
Params
−
+
−
+
−
+
−
+
Pre-Norm
58M
16
64
3.172
3.118
16
128 ‡
0.458
0.119
368M
4
20
2.759
2.660
8
50
1.069
0.360
1.3B
0.5
1
2.648
2.505
0.5
8 ‡
1.162
0.112
Peri-Norm
58M
32
64
3.152
3.120
32
128 ‡
0.440
0.087
368M
8
8
2.702
2.671
8
100 ‡
1.310
0.263
Appendix
Table 9: Learning-rate robustness in the dense sweeps (Figure 6 ), without ( − ) and with ( + ) LayerRoPE: the tuned rate η⋆ and loss L⋆ , the largest rate that trained stably, and LR sensitivity S ( Wortsman et al., 2024 ) , the mean of min(L(η),L0)−minηL(η) over the rates swept for both arms, with L0=ln32000 . Rates are in units of 10−3 . ‡ Largest rate swept.
Figure 18: Every loop depth at 770M. Validation loss throughout training for Parcae (solid) and Parcae + LayerRoPE ( dashed ), evaluated at T∈{1,2,4,8} loop iterations, from lighter to darker. The loss axis omits the first ∼1 B tokens.
Figure 19: Recurrent state during training. Mean norm of the recurrent state leaving the recurrent block (top, log scale) and mean pairwise cosine similarity between its centred token states (bottom), for Parcae and Parcae + LayerRoPE. Lines are rolling medians over logged steps, with raw values shown faintly.
Accuracy (%, ↑ )
Model
Looped Iters. (T)
Val. PPL ( ↓ )
Lambada PPL ( ↓ )
ARC-e
HellaS.-0
HellaS.-10
PIQA
LMB.
SQuAD
CoQA
QA-WD
Avg. ( ↑ )
140M [2pt]11.2 B
GPT
-
20.28
105.61
51.0
31.6
31.7
62.2
24.3
5.0
9.8
23.3
29.9
Parcae
8
18.66
81.17
53.6
33.7
33.7
63.1
26.4
7.5
12.8
33.1
33.0
Parcae + LayerRoPE
8
18.36
67.32
54.0
34.3
34.5
63.9
27.1
14.0
15.6
35.7
34.9
370M [2pt]29.6 B
GPT
-
15.01
37.16
59.5
40.3
40.1
67.0
32.9
20.2
20.1
44.4
40.6
Parcae
8
14.22
32.46
61.4
43.6
43.5
68.5
35.0
28.0
20.6
46.8
43.4
Appendix
Table 10: Per-task downstream results of the looped latent models. Validation perplexity on FineWeb-Edu, Lambada perplexity and accuracy (%) on each downstream task, averaged over three few-shot sampling seeds, with looped models at T=8 ; training tokens under each size, best per size in bold .
Figure 20: Learned depth multipliers at 1.3B. Learned magnitude multipliers er(ℓ) of LayerRoPE against the fixed 1/ℓ+1 of Layer-Norm Scaling.
Figure 21: Learned rotation angles at 1.3B. Rotation θ(ℓ) of the highest-frequency channel pair at each LayerRoPE site. Labels give the learned slopes βrot .
Figure 22: Effective depth gain at 1.3B. Magnitude ratio ∥γℓeff∥/∥γ0eff∥ of each layer’s effective normalization weights to the first layer’s, with every depth-dependent factor folded into γℓeff (legend). For Peri-Norm, the right column shows its output normalization; Pre-Norm and Layer-Norm Scaling have no residual-write weights. LayerRoPE’s weights γ are shared across layers, so its ratio is (ℓ+1)βmag ; labels give the learned slopes βmag .
Looped Transformers scale sequential computation by applying a compact stack of physical blocks for multiple rounds, increasing unrolled depth without increasing stored parameters. This reuse changes the residual-scaling problem: in an untied Transformer, each residual branch receives and applies its own parameter update, whereas in a looped Transformer one shared update aggregates gradients from repeated visits and is read back by those same visits in the next linearized forward pass. We formalize this tied-depth effect through a first-order perturbation bound controlled by a visit-alignment coefficient κR. The bound recovers the DeepNorm exponent when visits decorrelate, but in the conservative aligned regime it requires the exponent to increase from 1/4 to 1/2 as loop count grows at fixed physical depth. The resulting method, \textbf{DeepLoop}, keeps the Post-LN DeepNorm architecture and sets α=(2N)1/2 and β=(8N)−1/2 for unrolled depth N. On GPT-style looped language models at GPT-2 small and GPT-2 medium scale, DeepLoop is neutral when no physical block is revisited and improves validation loss and downstream accuracy once recurrent depth is activated. These results show that stable recurrent depth requires residual scaling rules that account for parameter visits, not only nominal layer count.
Shuzhen Li, Yifan Zhang, Jiacheng Guo +2
Princeton University · University of California, Los Angeles
Residual architectures are ubiquitous in deep learning, but they suffer from a subtle structural limitation: the norm of the residual stream can grow rapidly with depth. As a result, updates from later layers become small relative to the accumulated residual state. This reduces their impact on the representation and limits the benefits of scaling models in depth. To address this, we introduce NAG, a norm-agnostic residual architecture that separates magnitude from directional information in the residual stream, preserving meaningful layer contributions throughout depth and preventing later updates from being systematically suppressed by residual-norm growth. Importantly, NAG introduces only a negligible number of additional parameters and relies on simple operations that are easily kernel-fusible, preserving training efficiency in practice. We show that this architecture outperforms baseline Transformers, with gains that increase substantially as depth grows, enabling effective training of much deeper models. The norm-agnostic formulation also leads to an interpretable Mixture-of-Depths (MoD) mechanism that adaptively skips both attention and MLP layers. Beyond serving as a post-training accuracy-compute tradeoff, this mechanism can be used as a pretraining-time scaling strategy: under iso-FLOP training, compute saved by reducing per-token forward-pass cost can be reinvested into training on more tokens while keeping the total parameter count and KV-cache budget fixed. In our experiments, moderate Mixture-of-Depths rates of approximately 20%-25% match full-depth baseline performance under equal training compute while substantially reducing the number of executed layer parameters and forward-pass FLOPs. These results identify sparsity in depth as a new scaling axis for fixed-compute training, enabling very deep yet FLOP-efficient models.
Pre-norm is the standard normalization placement in modern Transformers because it facilitates joint optimization of full-depth models. We ask whether this preference persists when depth is introduced through a curriculum. In curriculum depth growth, each appended block receives the boundary representation produced by a trained prefix, making normalization placement relevant to forward conditioning. We therefore test whether placement and training curriculum interact. In a controlled distillation study with a Qwen3-8B teacher and a nine-layer student, pre-norm and post-norm are indistinguishable under joint training, differing by 0.0004 validation CE, while post-norm improves over pre-norm by 0.0328 under curriculum growth, an order of magnitude larger. A post-joint control matched by student active-layer tokens remains worse than post-grow, which rules out compute as the sole explanation. The ranking crosses over during the curriculum: post-norm takes the lead once blocks are appended. Single-block and freeze controls localize the ranking change to block appending rather than shallow-block quality or retraining. Boundary diagnostics associate post-norm with stable residual scales and pre-norm with structural-token scale drift; on a fixed batch, the final pre-grow block is also nearly identity-mapped. Together with the phase-wise crossover, these observations are consistent with boundary-scale conditioning after new blocks are appended. The results motivate treating normalization placement and training curriculum as coupled design choices in this distillation setting.
Sheng Ren, Yadong Wang, Naiqiang Tan +7
Nanjing University of Aeronautics and Astronautics, Nanjing, Jiangsu, China · Didichuxing Co. Ltd, Beijing, China