Post-training adapts language models in non-stationary environments. Practitioners monitor representation health with RankMe and related spectral statistics, often assuming that rank falls when representations degrade. We show that this assumption is unsafe for LLM post-training. In a controlled study of Qwen3-0.6B with four degradation modes and three seeds, data duplication worsens held-out loss by 75% relative to healthy while increasing both original and centred RankMe; the latter changes by 13.5 pooled standard deviations. Covariance effective rank rises to nearly twice its healthy value. This failure is spectral dispersion rather than collapse, so a one-sided monitor rates the worst checkpoint as the healthiest. By contrast, a learning-rate misconfiguration lowers centred RankMe and k95, while uncentred RankMe is inconsistent across seeds. Direction is therefore a property of the regime-statistic pair and cannot be fixed by recalibration alone. We also distinguish two often-conflated statistics: RankMe normalises singular values, whereas covariance effective rank normalises eigenvalues. On raw intermediate-layer states in the pretrained model, massive activations pin the latter near 1 out of dimension d while RankMe retains usable range. We then test a two-sided, multichannel sequential monitor with separate calibration and test data. In a pre-registered shared-prefix, leave-one-seed-out evaluation, it detects all three damage regimes in every fold 10 to 60 steps after the fork and separates dispersion from downward-rank damage by firing direction. However, it never precedes held-out probe loss, and calibration with two seeds produces false alarms on the held-out healthy seed. Spectral monitoring can diagnose failure regimes, but it does not warn earlier than held-out loss, and validity claims require held-out healthy data.
Figures & tables
transform
layer
reffΣ
reffΣ/d
RankMe
RankMec
λ1/T
none
14 (of 28)
1.06
0.001
120.13
121.24
0.995
none
last
192.90
0.188
770.63
822.50
0.134
RMS
14 (of 28)
308.96
0.302
718.24
782.25
0.058
RMS
last
192.19
0.188
765.11
818.28
0.131
Table 1: Effect of massive activations, and of per-token normalisation, on the statistics. Qwen3-0.6B, d=1024 ; the baseline here is the unmodified pretrained checkpoint (no fine-tuning, not a post-training endpoint), token-level; RankMe is the original uncentred definition, RankMec its centred variant. The covariance effective rank is floored on raw intermediate-layer states; neither form of RankMe is.
mode
probe loss
RankMec
reffΣ
k95
scb
healthy
2.530±0.01
812.6±0.4
174.8±1.6
689.7±0.6
0.394±0.00
high_lr
3.153±0.16
799.9±3.5↓ [3.6]
183.2±4.2 [1.9, n.s.]
665.7±5.9↓ [4.1]
0.277±0.01↓ [7.8]
duplicate_data
4.429±0.31
877.9±4.8↑ [13.5]
329.2±14.4↑ [10.7]
782.0±7.0↑ [13.1]
0.197±0.01↓ [17.6]
narrow_domain
2.586±0.01
816.4±0.8↑ [4.2]
179.4±1.1 [2.4]
696.3±1.5↑ [4.1]
0.368±0.00↓ [7.5]
Table 2: Endpoint diagnostics after 1000 steps, mean ± s.d. over 3 seeds; separation from healthy in pooled s.d. in brackets, defined as ∣μm−μh∣/σm2+σh2 over seed endpoints ( m = mode, h = healthy). Held-out probe loss is the pre-registered offline damage proxy. Arrows give the direction relative to healthy; bold marks a move opposite to the conventional “rank falls under degradation” reading.
Figure 1: The negative result at a glance (Qwen3-0.6B, shared-prefix branches). Left (seed 0): duplicate_data degrades held-out NLL most while its raw RankMe rises ∼ 100 above healthy — dispersion, not collapse. Right (all folds): loss breach ( × ) and ensemble alarm ( ∘ ) land within 10–60 steps of the fork and of each other — no consistent early warning — while the held-out healthy branch false-alarms at 60/460/450 steps (folds 0–2, as in Table 3 ).
high_lr
duplicate_data
domain_shift_v2
healthy (held-out)
pre-registered protocol (constant baseline)
ensemble alarm
10–20
20–30
20–60
FP@60/460/450
RankMec↓ only
10–20
cens.
20–30 (2/3)
FP@–/530/–
RankMec↑
550 (1/3)
20–30
120–380
FP@60/–/–
scb↓
30–40
40
50–60
FP@–/–/450
loss breach μL+3σL
10
20–30
10
none
Table 3: Shared-prefix, leave-one-seed-out detection (Qwen3-0.6B, last layer). Entries: steps after the fork, range over folds; “cens.” = no alarm in 600 steps; lead = breach − alarm step per fold (negative = alarm lags); last column: false alarms on the held-out healthy branch. Top: pre-registered; bottom: post-hoc fork-anchored, no pre-registration credit. RankMe rows: the protocol-frozen RankMec (uncentred: App. E ).