Different low-rank compression methods can produce compressed LLMs that respond differently to the same post-compression recovery procedure, and relative advantages observed between methods at the compression endpoint may shrink, grow, or even reverse after recovery. We ask whether this recovery heterogeneity reflects functional structure beyond scalar loss evolution, and how that structure evolves throughout recovery. Our results establish that this heterogeneity reflects a reproducible compression-induced functional structure, which we formalize as recovery pressure. To characterize this structure consistently throughout recovery, we develop a standardized functional characterization within each backbone that is applicable across heterogeneous low-rank methods. The primary backward characterization reveals reproducible module-wise structure across independent probes, while a complementary forward-only characterization recovers related structure without loss or backpropagation. We further find that recovery pressure measured at the endpoint is associated with subsequent recovery response; during recovery, its module-wise structure is reorganized non-uniformly, and localized updates induce distributed responses beyond directly updated modules. Further evidence indicates that tracking this evolving structure provides a complementary functional view of recovery progress alongside scalar loss.
Figures & tables
Figure 1 : Endpoint rankings need not be preserved after recovery. Avg7 MCQ Performance ( 3.4 ) on Qwen3-8B-Base, r=0.6 .
Figure 2 : Recovery reshapes comparative performance. Endpoint (hatched) and recovered (solid) Avg7 ( 3.4 ) across representative settings, with endpoint ranks annotated above each method.
Figure 3 : Fixed-coordinate recovery and backward characterization of recovery pressure. Recovery updates the trainable middle map while preserving compression-selected coordinates. Matched channel-wise perturbations at corresponding outputs of ct and d0 define the dense-referenced response pt , summarized by Pt and qm,t .
Figure 4 : Module-wise organization of endpoint recovery pressure. Colors show log10qm,0 for r=0.8 endpoints on the independent Q128 probe, normalized over attention and MLP modules.
Backbone
r
ρF,B median [range]
ρFsplit
Top-8
Llama-3.1 8B
0.8
.613 [.542, .910]
.987
5.5/8
0.6
.611 [.293, .832]
.983
5.0/8
Qwen3-8B Base
0.8
.377 [.042, .701]
.603
3.0/8
0.6
.471 [.295, .594]
.629
4.5/8
Table 1 : Forward-only grounding. Statistics summarize six compression methods within each backbone–retention setting. ρF,B is reported as median [range]; ρFsplit and Top-8 are medians.
Backbone
r
ρ(P0,HT)
Llama-3.1 8B
0.8
1.000
0.6
0.886
Qwen3-8B Base
0.8
0.771
0.6
0.829
Table 2 : Endpoint magnitude and recovery. Spearman ρ(P0,HT) across six low-rank methods.
Llama-3.1-8B r=.8
Llama-3.1-8B r=.6
Qwen3-8B-Base r=.8
Qwen3-8B-Base r=.6
Method
P0
Endpoint
Recovered
P0
Endpoint
Recovered
P0
Endpoint
Recovered
P0
Endpoint
Recovered
ASVD
1.314
33.55
55.14
1.285
32.55
44.86
0.241
59.61
62.40
1.464
34.60
52.35
SVD-LLM
0.192
51.94
58.80
0.567
36.81
50.92
0.336
55.93
60.57
0.430
43.12
54.57
Basis Sharing
0.374
51.17
57.97
1.425
33.67
48.47
0.450
53.97
60.42
0.680
44.03
53.23
MoDeGPT
0.074
61.10
62.45
0.159
45.94
53.80
0.421
54.48
59.35
0.550
41.50
49.77
GFWSVD
0.943
48.34
57.25
1.701
32.98
48.78
0.544
55.25
60.64
1.288
45.77
53.82
Table 3 : Recovery-aware comparison of low-rank compressed endpoints. P0 is measured on the Q128 probe before recovery. Endpoint and Recovered report zero-shot Avg7 (%) before and after 384 C4 updates under the shared fixed-coordinate intervention; recovered scores are averaged over three seeds. Bold marks the highest Endpoint or Recovered Avg7 within each backbone–retention setting. Full recovery gains and seed variation are reported in the Appendix A.2.9 .
Figure 5 : Nonuniform pressure reorganization under full and localized recovery. (a) Baseline-normalized module shares on Llama-3.1-8B ( r=0.8 , C4; mean ± range across 3 seeds). Flat line at 1.0 indicates uniform rescaling. (b) Share changes across unupdated modules after 64-step single-MLP recovery (SVD-LLM; 1.0 = no change).
Recovery updates
Avg7 ↑
C4 PPL ↓
Backbone
r
Loss
Loss+P
Update Red- uction(%)
Full
Loss
Loss+P
Full
Loss
Loss+P
Llama-3.1 8B
0.8
341.3
256.0
25.00
58.351
58.309
58.331
16.70
16.70
16.83
0.6
362.7
288.0
20.59
49.803
49.825
49.782
24.66
24.69
25.10
0.4
960.0
522.7
45.56
40.640
40.520
40.516
44.22
44.52
48.49
Qwen3-8B Base
0.8
330.7
341.3
-3.23
61.029
60.794
61.017
16.57
16.58
16.59
0.6
373.3
341.3
8.57
53.194
53.234
53.156
22.01
22.07
22.12
Table 4 : Controlled evaluation of pressure-informed stopping. Loss and Loss+P are evaluated on identical recovery trajectories and checkpoints, differing only by incorporating pressure dynamics. Each row averages six compression methods (Full uses 384 updates at r=0.8/0.6 and 1024 at r=0.4 ). Saving is relative to Loss; C4 PPL uses 256 validation windows of 256 tokens.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
r
Method
Llama-3.1-8B
Qwen3-8B-Base
Endpoint
+Recovery
Gain
P0
Endpoint
+Recovery
Gain
P0
0.8
ASVD
33.55 (6)
55.14 ±0.12 (6)
+21.60
1.544
59.61 (1)
62.40 ±0.62 (1)
+2.78
0.296
SVD-LLM
51.94 (2)
58.80 ±0.03 (3)
+6.86
0.699
55.93 (3)
60.57 ±0.92 (4)
+4.64
0.398
Basis Sharing
51.17 (4)
57.97 ±0.09 (4)
+6.79
0.868
53.97 (6)
60.42 ±0.49 (5)
+6.45
0.509
MoDeGPT
61.10 (1)
62.45 ±0.04 (1)
+1.35
0.662
54.48 (5)
59.35 ±0.35 (6)
+4.86
0.448
GFWSVD
48.34 (5)
57.25 ±0.23 (5)
+8.91
1.163
55.25 (4)
60.64 ±0.11 (3)
+5.39
0.609
Appendix
Table 5 : Model recoverability and endpoint pressure across low-rank baselines. Endpoint and +Recovery report zero-shot Avg7 (%) before and after 384-update BASIS recovery (mean ± SD across 3 seeds), with Gain showing net improvement. P0 measures the initial pressure magnitude on the C4 probe. Superscripts denote rank within each setting; bold marks the best score.
Llama-3.1-8B
Qwen3-8B-Base
Method
r=0.8
r=0.6
r=0.8
r=0.6
ASVD
19 / 10,725,120
21 / 11,137,728
31 / 10,963,770
20 / 11,261,960
SVD-LLM
14 / 10,958,080
19 / 11,151,936
13 / 11,191,752
17 / 10,975,608
Basis Sharing
18 / 10,892,160
25 / 11,083,200
17 / 11,122,488
22 / 10,794,960
MoDeGPT
4 / 10,878,704
5–6 / 10,994,530
4 / 10,826,464
5–6 / 10,990,644
GFWSVD
14 / 10,958,080
19 / 11,151,936
13 / 11,191,752
17 / 10,975,608
Appendix
Table 6: Native BASIS capacity for the 384-update C4 recovery. Each cell gives adapter rank ra / exact trainable parameter count for one frozen compression endpoint. Counts are identical across the three recovery seeds. Shared Basis Sharing adapters are counted once per shared basis.
A. Within-setting endpoint rank associations
Probe
Backbone
r
ρ(P0,L0)
ρ(P0,H384)
ρ(L0,H384)
Q0
Llama
0.8
1.000
1.000
1.000
Q0
Llama
0.6
0.886
0.886
1.000
Q0
Qwen
0.8
0.771
0.771
1.000
Q0
Qwen
0.6
0.943
0.829
0.943
Q128/V128
Llama
0.8
1.000
1.000
1.000
Appendix
Table 7 : Endpoint magnitude, scalar endpoint quality, and subsequent recovery. Panel A reports within-setting Spearman correlations across the six compression methods. L0 is held-out endpoint NLL and H384=L0−L384 is the mean held-out loss reduction over three recovery seeds. Q0 denotes the original 16-window pressure probe; Q128/V128 independently remeasures all endpoints using separate pressure and validation sets. Panel B reports a retrospective method-held-out prediction audit over all 24 Q128/V128 endpoints.
Figure 6 : Factor-update recovery shifts selected coordinates. Collapse-score ECDFs on Llama-3.1-8B show that directly updated factors in U-LoRA, V-LoRA, and UV-LoRA move away from the whitening reference, while BASIS keeps the protected outer factors fixed.
Recovery method
Endpoint Avg7 ↑
Avg7 projection diagnostic
PPL Avg ↓
Learned
Projected
Total
Projected gain
Coord.-shift
Shift share
Learned
Projected
BASIS
50.48
50.48
+17.29
+17.29
+0.00
0.0%
16.57
16.57
U-LoRA
50.15
39.17
+16.96
+5.98
+10.98
64.7%
17.34
223.96
V-LoRA
49.91
41.80
+16.73
+8.61
+8.12
48.5%
17.90
119.25
JFT
50.26
40.85
+17.07
+7.67
+9.40
55.1%
17.56
283.05
Appendix
Table 8: Controlled fixed-coordinate projection diagnostic for factor-update recovery (LLaMA-3.1-8B, r=0.4 ). Normal evaluates each method in its learned coordinates; Projected maps recovered back to Source-defined fixed-coordinate subspace.
Llama-3.1-8B
Qwen3-8B-Base
Method
Spearman
Cosine
Spearman
Cosine
ASVD
0.996
0.963
0.993
0.996
SVD-LLM
0.968
0.983
0.978
0.997
Basis Sharing
0.988
0.996
0.991
0.999
MoDeGPT
0.358
0.337
0.493
0.997
GFWSVD
0.991
>0.999
0.988
>0.999
Appendix
Table 9: Reproducibility of endpoint module organization. Spearman rank correlation and cosine similarity compare module-share profiles from Q64 and independent Q128 probes over all attention and MLP modules.
Llama-3.1-8B
Qwen3-8B-Base
Method
Spearman
Cosine
Spearman
Cosine
ASVD
0.958
0.477
0.948
0.986
SVD-LLM
0.913
0.739
0.936
0.999
Basis Sharing
0.960
0.905
0.948
0.998
MoDeGPT
0.201
0.338
0.539
0.986
GFWSVD
0.968
0.983
0.907
>0.999
Appendix
Table 10: Module-profile sensitivity to the original probe size. Module-share agreement between the original 16-window Q0 and independent Q128 probes at r=0.8 .
Llama-3.1-8B
Qwen3-8B-Base
Method
Spearman
Cosine
Top-8
Spearman
Cosine
Top-8
ASVD
0.940 ∗
0.971
5/8
0.662
0.783
7/8
SVD-LLM
0.614
0.690
6/8
0.811
0.813
5/8
Basis Sharing
0.738
0.983
7/8
0.877
0.822
6/8
MoDeGPT
0.934
0.992
7/8
0.714
0.481
3/8
GFWSVD
0.869
0.955
6/8
0.848
0.592
5/8
Appendix
Table 11: Forward–backward characterization alignment. Module-wise agreement on matched two-window inputs at r=0.8 . This is a separate small-input/module-level audit and is not the source data summarized in Table 1.
Figure 7 : Absolute-energy responses to localized recovery. Normalized energy change ( Em,64/Em,0 on Q128; 3-seed mean) across 60 unupdated modules under single-MLP BASIS recovery (SVD-LLM, Llama-3.1-8B, r=0.8 ). Rows denote updated layers; blue indicates attenuation ( <1.0 ) and gray marks excluded sites. Relative share growth can coexist with absolute energy decay.
Method
Seed
TV
Ref.
E↓
SVD-LLM
7711
0.479
0.210
64
SVD-LLM
7712
0.414
0.290
64
SVD-LLM
7713
0.369
0.311
64
Swift-SVD
7711
0.334
0.255
64
Swift-SVD
7712
0.436
0.290
64
Swift-SVD
7713
0.376
0.276
64
Appendix
Table 12: Full-profile change over steps 0–384 for each selected recovery seed. Ref. is the paired-document error reference. E↓ counts absolute-energy decreases out of 64 modules; these sign counts are descriptive.
Method
64
128
192
256
320
384
SVD-LLM
0.355/0.130
0.443/0.154
0.398/0.177
0.366/0.219
0.415/0.148
0.417/0.191
Swift-SVD
0.366/0.175
0.410/0.122
0.401/0.227
0.318/0.184
0.365/0.199
0.379/0.254
MoDeGPT
0.312/0.217
0.336/0.183
0.487/0.237
0.354/0.258
0.311/0.299
0.370/0.243
Appendix
Table 13: All measured checkpoints: seed-mean profile TV and its paired-document error reference. Cells contain TV / Ref.
Backbone
Method
TV
Ref.
Llama
ASVD
0.493
0.271
Llama
SVD-LLM
0.479
0.210
Llama
Basis Sharing
0.495
0.293
Llama
MoDeGPT
0.414
0.288
Llama
GFWSVD
0.768
0.265
Llama
Swift-SVD
0.334
0.257
Appendix
Table 14: Complete endpoint cohort at retention 0.8; one recovery seed.
Method
L1 A
L16 A
L31 A
L1 M
L16 M
L31 M
SVD-LLM
0.489
0.124
5.307
69.276
0.204
6.368
Swift-SVD
0.518
0.125
5.214
69.105
0.204
6.257
MoDeGPT
0.265
0.131
24.457
44.557
0.206
16.196
Appendix
Table 15: Initial module shares (percent) for the layers in the main figure. A and M denote attention and MLP. These baselines are shared across recovery seeds.
Recovery updates
Avg7 ↑
Backbone
r
Seed
Loss
Loss+P
Update reduction (%)
Full
Loss
Loss+P
Δ Avg7
Llama-3.1-8B
.6
7711
362.7
288.0
20.59
49.803
49.825
49.782
−0.043
7712
362.7
288.0
20.59
49.074
49.241
49.248
+0.007
7713
384.0
320.0
16.67
50.118
50.118
49.648
−0.470
Mean
369.8
298.7
19.23
49.665
49.728
49.559
−0.169
Llama-3.1-8B
.4
7711
960.0
522.7
45.56
40.640
40.520
40.516
−0.004
Appendix
Table 16 : Seed-level controlled stopping comparison. Each seed row averages six compression methods; mean rows aggregate all 18 method–seed trajectories within each setting. Full recovery uses 384 updates at r=.6 and 1024 at r=.4 . Update reduction is measured relative to Loss ; Δ Avg7 is Loss+P minus Loss .
Backbone
r
Scale s
Loss
Loss+P
Update reduction (%)
Δ Avg7 (pp)
Within ±1 pp
Llama-3.1-8B
.6
0.5×
362.7
352.0
2.94
−0.112
6/6
1.0×
362.7
288.0
20.59
−0.043
6/6
1.5×
352.0
266.7
24.24
+0.047
6/6
2.0×
341.3
245.3
28.12
−0.012
5/6
.4
0.5×
970.7
661.3
31.87
−0.003
6/6
1.0×
960.0
522.7
45.56
−0.004
6/6
Appendix
Table 17 : Controlled stopping comparison across nearby loss operating points. The loss threshold is scaled by s∈{0.5,1.0,1.5,2.0} while the remaining stopping protocol is fixed. Loss and Loss+P report mean recovery updates across six compression methods. Update reduction is measured relative to Loss ; Δ Avg7 is Loss+P minus Loss in percentage points. The final column reports the number of methods whose Avg7 difference remains within ±1 pp. We use seed 7711 that is also in Table 4