Persistent Depth Ordering amid Shifting Block-Bypass Responses in Language Model Pretraining
Authors: Shengye Tao, Yinzhu Cheng, Haihua Xie
Organizations: Beijing University of Civil Engineering and Architecture · Beijing Institute of Mathematical Sciences and Applications (BIMSA) · Institute of Statistics and Big Data, Renmin University of China
Layer interventions are widely used to probe the internal organization of language models, yet most analyses examine a single training checkpoint even though model representations and computations evolve throughout pretraining. This leaves open which depth-dependent intervention responses reflect persistent organization and which are transient consequences of training. We study this question using single-block identity bypass on fixed teacher-forced contexts across five released trajectories and 11 model-domain combinations. We find that block-bypass responses retain recognizable depth ordering while their magnitudes redistribute: nearby checkpoints preserve stronger rank correspondence than distant ones, and large changes concentrate at positions that recur across text samples and transfer across evaluation domains. Controlled experiments further show that changes in the natural bypass effect cannot be reduced to a single downstream sensitivity: in replicated Pythia runs, local missing-update magnitude grows while the pooled matched downstream response decreases, whereas OLMo-2 7B exhibits a different balance. These matched responses also depend on perturbation strength and direction, without identifying targeted compensation. Together, our results show that longitudinal layer sensitivity is structured but not static, and that single-checkpoint intervention responses should be interpreted in the context of how the underlying perturbation pathway evolves during training.
Figures & tables
Figure 1: Measurement and paper roadmap. (a) Fixed teacher-forced contexts pass through an intact block or identity bypass, then the same suffix and output head. We follow response fields over training, then examine continuous outputs, matched perturbations and direction controls. (b) Historical Pythia-160M Hella responses average six runs over ten interior blocks at 14k, 72k and 143k; entries are percentages. (c–e) Means, ascending ranks and bottom-two selections derive from the same 30 mean-profile cells. Block indices are zero-based. Columns are observed checkpoints, not interpolated intervals.
Model
Measured blocks
Checkpoints
Text
Runs
OLMo-2 1B
14
5
G/M/C
1
OLMo-2 7B
30
5
G/M/C
1
OLMo-2 13B
38
5
G
1
SmolLM2 1.7B
22
5
G/M/C
1
Pythia 1.4B
22
5
G
1
Pythia 160M
10
3
Hella-derived
6
Table 1: Model and text coverage. G/M/C denote general, mathematical, and code text. The released models each contribute one training trajectory. Pythia-160M provides repeated training runs for the controlled experiments.
Figure 2: Depth ordering persists as response magnitudes evolve. (a,b) General-text profiles at the first, middle, and last observed checkpoints. Depth is normalized over the measured interior blocks. (c) Thin lines show each model–domain combination’s median correlation at a given checkpoint separation; the heavy line summarizes all pairs. Separations 1–4 contain 44, 33, 22, and 11 pairs.
Figure 3: Response changes concentrate at repeatable positions. (a) Cumulative absolute endpoint change, with blocks ordered by change magnitude. Thin curves show model–domain combinations; the dark curve is their pointwise median. The diagonal represents equal change at every block. (b) The largest-changing positions are selected on nine windows and evaluated on the other nine. Points and bars show medians and 5th–95th percentiles over 1,000 partitions. Gray ticks show the random-selection expectation. O1/S denote OLMo-2 1B/SmolLM2; G/M/C denote the three text domains.
Figure 4: A trajectory’s redistribution pattern transfers across text domains. (a–c) Endpoint changes in percentage points, with model-specific vertical scales. Blue, red, and green denote general, mathematical, and code text. (d) Each domain is correlated with a standardized template constructed from the other two. Filled points include 95% paired-window-bootstrap intervals; hollow points show estimates with unavailable window-level uncertainty. The template and held-out domain use the same training endpoints.
Figure 5: The training effect extends to continuous output distributions. (a) Pythia disagreement in six paired runs, with their mean shown by the heavy line. (b) JS increases in every intact-margin stratum, using fixed early-checkpoint cutpoints. (c) Run-averaged JS profiles retain depth ordering while increasing in magnitude ( ρ=0.915 ). Blue and red denote early and final checkpoints.
Figure 6: Training changes the balance between local perturbation and downstream response. (a) Natural local relative RMS. (b) JS at matched relative magnitude, adjusted for the perturbation after FP16 rounding and pooled over four scales. Thin Pythia lines show runs; heavy lines show their means. OLMo contributes one trajectory. (c) Pythia checkpoint changes at each target scale, with paired-run intervals.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Panel
ΔDˉ (pp)
ρ
Affine R2
Ck (%)
High
Low
OLMo-2 1B / G
+2.37
0.675
0.461
54.5
0.67
1.00
OLMo-2 1B / C
+3.77
0.741
0.570
54.6
1.00
0.67
OLMo-2 1B / M
+2.65
0.380
0.411
54.9
0.33
0.33
OLMo-2 13B / G
+1.89
0.716
0.622
39.6
0.75
0.38
OLMo-2 7B / G
-1.46
0.785
0.543
67.3
1.00
0.50
OLMo-2 7B / C
-1.71
0.628
0.303
65.5
0.83
0.50
Appendix
Table 2: Endpoint response, redistribution, and identity retention. High/Low are the overlap fractions of the highest/lowest response sets. G/M/C denote general, mathematical, and code text. The affine fit measures numerical correspondence after a common shift and scale.
Panel
ρˉadj−ρend
95% interval
Share (%)
5–95% splits
OLMo-2 1B / G
0.171
[0.091, 0.323]
51.0
[48.6, 53.9]
OLMo-2 1B / C
0.201
[0.107, 0.305]
54.0
[52.9, 54.7]
OLMo-2 1B / M
0.402
[0.176, 0.423]
51.7
[47.2, 54.7]
SmolLM2 / G
0.283
[0.202, 0.384]
84.2
[83.4, 84.7]
SmolLM2 / C
0.231
[0.172, 0.315]
86.8
[85.3, 87.5]
SmolLM2 / M
0.298
[0.207, 0.402]
76.6
[75.6, 76.8]
Appendix
Table 3: Window-level evidence for persistence and localized change. Correlation-gap intervals use paired-window bootstrapping. The remaining columns show held-out change shares and their partition percentiles.
Panel
Shallow (%)
Middle (%)
Deep (%)
Shallow, trimmed (%)
OLMo-2 1B / G
58.7
20.6
20.7
45.8
OLMo-2 1B / C
54.8
9.6
35.5
45.2
OLMo-2 1B / M
63.1
15.2
21.7
52.5
OLMo-2 13B / G
31.3
41.4
27.4
24.0
OLMo-2 7B / G
70.0
14.4
15.5
60.3
OLMo-2 7B / C
69.2
16.5
14.3
53.5
Appendix
Table 4: Where endpoint changes accumulate. The trimmed column gives the shallow share after excluding the first measured interior block. Regional boundaries are unchanged.
Figure 7: A common shift and scale leave depth-specific residuals. SmolLM2/code and Pythia-1.4B/general illustrate the lowest and highest affine R2 values. Red shows the observed final profile; blue shows the fitted transformation of the first profile. Shading marks their difference.
Panel
High: within
High: across
Low: within
Low: across
OLMo-2 1B / G
0.83
0.67
0.83
0.83
OLMo-2 1B / C
0.83
0.83
0.83
0.67
OLMo-2 1B / M
0.83
0.33
0.67
0.33
SmolLM2 / G
0.90
0.80
0.80
0.30
SmolLM2 / C
0.80
0.60
0.70
0.60
SmolLM2 / M
0.80
0.50
0.90
0.60
Appendix
Table 5: Block-set stability with matched text budgets. Entries are medians over 1,000 half-window partitions. High/Low use the same set sizes as Table 2 . Within-checkpoint overlap provides a text-sampling comparison for across-checkpoint retention.
Trajectory
Held-out text
All measured
Trimmed ends
OLMo-2 1B
Code
0.875
0.881
OLMo-2 1B
General
0.933
0.935
OLMo-2 1B
Math
0.944
0.918
OLMo-2 7B
Code
0.934
0.920
OLMo-2 7B
General
0.944
0.962
OLMo-2 7B
Math
0.986
0.959
Appendix
Table 6: Leave-one-domain-out redistribution patterns. All comparisons hold the trajectory and training endpoints fixed. Trimming removes the first and last measured blocks before constructing the template.
Figure 8: Successive sensitivity changes across text domains. G/M/C denote general, mathematical, and code text. Rows show adjacent-checkpoint differences, and columns show measured interior blocks. Red indicates an increase and blue a decrease. Each model has its own symmetric color scale. The recurring depth patterns coexist with differences in change magnitude.
Endpoint
Pythia change
95% run interval
OLMo change
Top-1 disagreement
+0.0792
[0.0677, 0.0881]
-0.0146
Forward KL
+0.2492
[0.2194, 0.2722]
-0.0269
Jensen–Shannon
+0.0493
[0.0435, 0.0537]
-0.0053
Target NLL change
+0.2498
[0.2234, 0.2710]
-0.0313
Logit-margin loss
+0.1066
[0.0792, 0.1292]
-0.0274
Appendix
Table 7: Continuous and discrete response changes. Pythia intervals resample six paired training runs. OLMo contributes one trajectory. Disagreement is in probability units; multiplying by 100 gives percentage points.
Endpoint
Pythia ρ
OLMo ρ
Top-1 disagreement
0.903
0.785
Forward KL
0.891
0.787
Jensen–Shannon
0.915
0.766
Target NLL change
0.879
0.705
Logit-margin loss
0.430
0.547
Appendix
Table 8: Depth-rank persistence across output metrics. Values are early–final correlations of the continuous-output experiment’s layer profiles.
γ
Matched-JS change
95% run interval
0.25
+0.001291
[+0.000866, +0.001784]
0.50
+0.000114
[-0.000739, +0.000945]
0.75
-0.003289
[-0.005309, -0.001642]
1.00
-0.008606
[-0.014081, -0.004576]
Appendix
Table 9: Pythia matched-JS changes by target scale. The primary adjusted comparison pools the four scales; these scale-specific comparisons reveal their different checkpoint trends.
Figure 9: Downstream perturbation changes accumulate late in the suffix. Final-minus-early contrasts in hidden-state relative RMS at injection, successive suffix regions, and final normalization. Pythia intervals resample paired training runs; OLMo contributes one trajectory. The increasingly negative Pythia contrast indicates a larger checkpoint difference near the output boundary.
Predictors
Pythia test R2
OLMo test R2
Local magnitude
0.514
0.394
Downstream propagation
0.015
0.074
Magnitude + propagation
0.543
0.457
Joint + checkpoint terms
0.623
0.456
Appendix
Table 10: Held-out prediction of log JS. Models use local magnitude, downstream propagation, or both. The final model adds checkpoint-dependent terms. OLMo scores describe one trajectory.
Figure 10: Directional specificity varies with depth and control family. Early and final Pythia contrasts use paired-run intervals. Negative values indicate stronger attenuation of the natural direction. Random-orthogonal and token-permuted controls give different depth patterns. Neither control comparison establishes targeted compensation.
Contrast
Early
Final
Change
95% change interval
Natural − random
+0.0966
-0.0328
-0.1294
[-0.1450, -0.1125]
Natural − permuted
+0.0406
+0.0276
-0.0130
[-0.0318, +0.0066]
Appendix
Table 11: Adjusted Pythia directional specificity. Checkpoint changes are paired within the same runs and control constructions.
Added information
Test R2 increment
95% run interval
Readout projection
0.00232
[0.00189, 0.00277]
Generic propagation
0.10146
[0.09061, 0.11245]
Direction-family identity
0.00783
[0.00606, 0.00915]
Appendix
Table 12: Held-out information from directional controls. Intervals resample the six Pythia runs. The preregistered threshold for a material increment is 0.01 in R2 .
Figure 11: Historical Pythia replication at run, layer and checkpoint resolution. (a,b) All 330 run–block–checkpoint cells use a common 0–70% scale. Asterisks identify shared runs 6 and 8. (c) Open points show 110 run–block endpoint pairs; filled points are run-averaged layers. (d,e) Thin traces retain runs and bold traces show panel means. (f) All six Hella and five corrected WikiText-2 run-mean changes are positive, averaging +7.93 and +9.29 pp. The Hella whisker is the retained 95% paired-run interval [6.77,8.85] pp; no WikiText interval is introduced. Layers and repeated displays are not independent training replications.
Figure 12: Five released general-text response fields. (a) Mean changes share vertical limits and use recorded within-stage progress; initial absolute means are printed above each model. (b) All 630 interior-block/checkpoint responses use a common 0–100% scale. Relative depth is block index divided by total blocks minus one. Narrow strips denote the five observed samples; gaps are unmeasured. (c) Endpoint stems are descriptive point estimates, not confidence intervals. OLMo-2 7B ends lower despite partial recovery; OLMo progress refers to stage one. Each model supplies one trajectory.
When a language model fails on surface-perturbed input (typos, OCR noise, homophones), "which layer is responsible" has three natural operationalizations: where representations diverge most (sensitivity), where restoring clean activations recovers the prediction (causality), and where a small adapter can repair the damage (compensatory capacity) - and we show these three layer maps dissociate. Across a five-model panel we identify two propagation regimes - spike-and-suppress (Phi-3.5, Gemma-2-9B) and late-accumulation (Llama-3, Mistral, Qwen2.5-7B) - and on the two models meeting an 80% identity-patch gate, sensitivity and causality are anti-correlated (rho = -0.72 to -0.88). Within-family scaling on Qwen2.5 (1.5B to 14B) shows the late-accumulation signature strengthening monotonically with scale, corroborated on a second family. We propose cascade disruption as the mechanism behind the dissociation: adapters placed at causally implicated early layers break intact downstream computation, making diagnostic-flagged sites the worst adapter placements. A fixed-harness layer sweep across four models (3.8-8B) confirms the core prediction on chain-of-thought GSM8K - the flagged sites are the most damaging adapter windows on every adjudicable model - and is sign-consistent but strongly attenuated on a multiple-choice control, consistent with damage that compounds with generation length. The sweep yields practical guidance: a training-free LRD pre-screen and a default-deepest placement rule, though absolute gains over no-adapter baselines remain small. Finally, apparent gains from a representation-stability loss reverse under an adequate generation budget - truncated chain-of-thought had been scored as empty - a methodological warning for any intervention evaluated on chain-of-thought tasks.
A dense, pretrained language model can be retrofitted with recurrent depth and learn an iterative latent transition that persists after outcome-only annealing. Qwen2.5-0.5B-Instruct is split into a Prelude, a weight-tied Recurrent Block, and a Coda, with an identity-preserving one-loop path and a re-entry bridge on later loops. At loop 1 the retrofit remains non-inferior to its base on a preregistered ARC battery. Three findings. First, the mechanism is a reusable procedure rather than terminal-answer lookup, and installs at two budgets: 6M trained parameters over frozen base weights and 180M full-block. With intermediate-step supervision, the model computes one task step per loop and persists when only final answers are graded. The adapter matched the full block overall (83.8% versus 84.0%), led through depth 11, and trailed beyond. Verbal fine-tuning reached 79-86% on controlled verbal renderings (zero-shot transfer was minimal), and adapter verbal training begun from the installed mechanism outpaced matched fresh training by 18.6 points, including on a held-out test set. Second, the operation extrapolates to roughly 1.5 times its supervised depth, holding 70% accuracy through depth 18. Third, a same-size scratchpad-trained model matched the recurrent model within its learned horizon but collapsed beyond it. The recurrent model won overall, 84% versus 72%, retained 53% versus 2.5% beyond depth 10, and answered 7.6 times faster. An iterative transformer can therefore perform deeper reasoning in latent space faster than comparable or larger models fine-tuned on the same task, in a system-level comparison. A second task, running the rule in reverse, exposed the limits: the inverse was learnable in isolation, but no continuation acquired it while preserving the installed mechanism and general capability, a catastrophic-interference boundary. Learned depth selection remains open.
A causal-decoder block is hierarchical: lower layers build the residual basis that upper layers attend over. We identify a failure mode in GPT pretraining: upper layers commit to sharp attention patterns before lower-layer features stabilize. We call this premature upper-layer attention specialization. Temporarily slowing only upper-layer Q/K projections during early training improves final perplexity and downstream accuracy without altering other parameters; it prevents upper attention from collapsing onto an immature residual basis. In LLaMA-style blocks, the same intervention is nearly unnecessary. Through ablations, we isolate multiplicative gated FFNs (not RMSNorm or bias removal) as the component that suppresses the upstream residual writes driving the failure. A pathwise analysis unifies both findings: the learning-rate intervention reduces a step-size factor, while gated FFNs reduce a residual-energy factor on the same growth pathway. Our results identify upper-layer Q/K timing as a concrete interaction point between decoder architecture and optimization.
Jinchang Zhu, Jindong Li, Yuwen Hao +3
The Hong Kong University of Science and Technology (Guangzhou)