Frozen Factor or Spectral Band? Disentangling Two Choices in Low-Rank LoRA
Organizations: CentraleSupélec, Université Paris-Saclay · École polytechnique, Institut Polytechnique de Paris
Abstract
Spectral variants of low-rank adaptation (LoRA) choose both a subspace and which factor to freeze. We separate these choices by freezing the input factor A or output factor B on the top or bottom singular directions of pretrained weights, with learning rates selected separately. At rank 2, the same-band advantage of freezing A is larger than either within-factor band difference on all four task-model pairs with complete comparisons. Freezing B also trails comparable-budget free LoRA by 8-18 percentage points on five pairs spanning a formatting task and OpenBookQA. The A-frozen advantage persists in a single-GPU-model replication and within individual MLP module groups, including controls with equal or greater trainable counts for B frozen, and when A is frozen on a random orthonormal basis. The factor contrast weakens with rank. On OpenBookQA / Qwen2.5-1.5B at rank 16, PEFT's MiCA implementation trails comparable-budget LoRA by 3.08 points under a shared training recipe transferred from the MiCA paper. A trained oracle output subspace largely removes the low-rank deficit; partial warm-up gains recur across three direction seeds. The factor-versus-band ordering is descriptive; an approximate multiplicity audit weakens several earlier significance claims. These results extend known factor asymmetry by showing how its magnitude depends on spectral placement, rank and training conditions.
Figures & tables
| Configuration | Initialization and constraint | Trainable parameters/module |
|---|---|---|
| B frozen, band | fixed; train from zero | |
| A frozen, band | fixed; train from zero | |
| Both trainable, band | Initialize and ; train both | |
| Free LoRA | Kaiming-initialized , zero ; train both |
| Task / model | Factor | Band | Interaction |
|---|---|---|---|
| Format / Llama-3.2-1B | |||
| OpenBookQA / Qwen2.5-1.5B | |||
| OpenBookQA / Qwen2.5-7B | |||
| Model | Band | Extra count | A( ) B( ) | |
|---|---|---|---|---|
| Qwen2.5-1.5B | Top | |||
| Bottom | ||||
| Qwen2.5-7B | Top | |||
| Bottom |
| Task / model | Adapted modules | A rank / B rank | Top A B | Bottom A B |
|---|---|---|---|---|
| Format / Llama-1B | Down | 2 / 2 | ||
| Gate + up | 2 / 8 | |||
| OBQA / Qwen-1.5B | Down | 2 / 2 | ||
| Gate + up | 2 / 12 |
| A frozen minus B frozen | Free LoRA minus B frozen | |||||
|---|---|---|---|---|---|---|
| Frozen rank | Top | Bottom | Top | Bottom | ||
| 2 | ||||||
| 4 | ||||||
| 8 | ||||||
| 16 | ||||||
| Configuration | Training recipe | Accuracy (%) | Supervised tokens | Steps |
|---|---|---|---|---|
| Bottom B frozen, | Main | 49.84 | 602,112 | 49 |
| Free LoRA, | Main | 56.12 | 602,112 | 49 |
| PEFT MiCA, | Transferred | 52.40 | 417,792 | 272 |
| PEFT free LoRA, | Transferred | 55.48 | 417,792 | 272 |
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| B frozen, | A frozen, | Free | Base | |||
|---|---|---|---|---|---|---|
| Task / model | Top | Bottom | Top | Bottom | ||
| Format / Llama-3.2-1B | 64.43 | 59.22 | 77.39 | 78.18 | 76.84 | 0.00 |
| Format / Qwen2.5-1.5B | 62.57 | 63.91 | — | — | 78.79 | 0.00 |
| OpenBookQA / Qwen2.5-1.5B | 46.52 | 42.20 | 53.08 | 52.48 | 54.56 | 42.20 |
| OpenBookQA / Qwen2.5-7B | 53.16 | 48.40 | 63.64 | 61.48 | 62.84 | 47.40 |
| OpenBookQA / Mistral-7B | 47.64 | 47.76 | 60.76 | 58.60 | 60.12 | 44.60 |
| Both trainable, | Both trainable, | |||
|---|---|---|---|---|
| Task / model | Top init. | Bottom init. | Top init. | Bottom init. |
| Format / Llama-3.2-1B | 79.28 | 79.54 | 80.00 | 80.72 |
| Format / Qwen2.5-1.5B | 78.96 | 78.11 | 77.75 | 78.57 |
| OpenBookQA / Qwen2.5-1.5B | 55.24 | 54.96 | 55.76 | 55.28 |
| OpenBookQA / Qwen2.5-7B | 62.32 | 62.64 | 62.92 | 63.12 |
| OpenBookQA / Mistral-7B | 59.24 | 59.24 | 59.88 | 60.48 |
| B frozen | A frozen | Free | Ref. rank | ||||
|---|---|---|---|---|---|---|---|
| Task / model | Rank | Top | Bottom | Top | Bottom | ||
| OBQA / Qwen2.5-1.5B | 4 | 46.44 | 45.24 | 53.44 | 53.88 | 55.24 | 2 |
| 8 | 49.92 | 46.56 | 56.28 | 53.44 | 55.44 | 4 | |
| 16 | 54.08 | 49.84 | 55.44 | 53.64 | 56.16 | 7 | |
| OBQA / Qwen2.5-7B | 16 | 60.20 | 56.76 | 62.40 | 62.56 | 63.20 | 7 |
| ARC-C / Qwen2.5-1.5B | 2 | 48.42 | 46.68 | — | — | 50.44 | 1 |
| Model / task | Rank | Top A B | Bottom A B | Interpretation |
|---|---|---|---|---|
| Qwen2.5-1.5B / OBQA | 4 | Both pass | ||
| 8 | Both pass | |||
| Qwen2.5-7B / OBQA | 16 | Both pass |
| Task / model, frozen rank 16 | B top | B bottom | Free reference | Free minus bottom |
|---|---|---|---|---|
| Format / Llama-3.2-1B | 77.56 | 75.54 | 82.05 ( ) | 6.51 |
| Format / Qwen2.5-1.5B | 74.30 | 75.24 | 80.78 ( ) | 5.54 |
| OpenBookQA / Qwen2.5-1.5B | 54.08 | 49.84 | 56.16 ( ) | 6.32 |
| OpenBookQA / Qwen2.5-7B | 60.20 | 56.76 | 63.20 ( ) | 6.44 |
| ARC-Easy / Qwen2.5-1.5B | 79.48 | 77.92 | 79.40 ( ) | 1.48 |
| ARC-Challenge / Qwen2.5-1.5B | 49.70 | 44.52 | 49.52 ( ) | 5.00 |
| B-frozen configuration | Earlier rate | Revised rate | Earlier score | Revised score |
|---|---|---|---|---|
| Top, rank 1 | 51.07 | 48.34 | ||
| Bottom, rank 1 | 45.47 | 45.47 | ||
| Random spectral indices, rank 1 | 52.88 | 52.77 | ||
| Random spectral indices, rank 4 | 70.23 | 70.81 |
| Control and contrast | Top | Bottom | Scope |
|---|---|---|---|
| Fresh seeds: both trainable minus B frozen | 14.50 | 21.27 | Format / Llama; seeds 120–124 |
| Four-times tokens: same contrast | 11.99 | 18.24 | Format / Llama; all rates retuned |
| Original-token counterpart | 15.57 | 21.50 | Same task, original campaign |
| Answer-only: A frozen minus B frozen, | 7.16 | 9.72 | OBQA / Qwen-1.5B; steps fixed |
| Square-module control: A frozen minus B frozen | 39.3 | 33.6 | Top arm meets collapse condition |
| Configuration | Rank | Selected rate | Original score | Replication |
|---|---|---|---|---|
| B frozen, top | 2 | 64.43 | 64.23 | |
| B frozen, bottom | 2 | 59.22 | 58.31 | |
| A frozen, top | 2 | 77.39 | 77.39 | |
| A frozen, bottom | 2 | 78.18 | 77.98 | |
| Free LoRA | 1 | 76.84 | 76.97 | |
| Both trainable, top initialization | 1 | 79.28 | 80.36 |
| Pair | Group | Configuration | Score top / bottom | Rate top / bottom |
|---|---|---|---|---|
| Format / Llama-1B | Down | B frozen, | 29.38 / 12.30 | / |
| A frozen, | 65.92 / 67.75 | / | ||
| Free, | 73.49 | |||
| Gate + up | B frozen, | 39.15 / 40.53 | / | |
| A frozen, | 78.14 / 75.53 | / | ||
| B frozen, | 65.57 / 65.49 | / |
| Configuration | Selected rate | Answer-only score | Full-loss score |
|---|---|---|---|
| B frozen, top | 48.36 | 46.52 | |
| B frozen, bottom | 45.76 | 42.20 | |
| A frozen, top | 55.52 | 53.08 | |
| A frozen, bottom | 55.48 | 52.48 | |
| Free LoRA, rank 1 | 55.72 | 54.56 |
| Task, rank | Rate | Seed 2 | Seed 3 | Seed 4 | Seed 5 | Seed 6 | Mean |
|---|---|---|---|---|---|---|---|
| Format, 2 | 77.52 | 78.18 | 77.69 | 76.71 | 77.04 | 77.43 | |
| OpenBookQA, 2 | 55.20 | 55.00 | 53.80 | 55.60 | 53.20 | 54.56 | |
| OpenBookQA, 16 | 56.20 | 56.20 | 56.20 | 55.80 | 57.80 | 56.44 |
| Task, frozen rank | Minus A-top | Minus A-bottom | Minus free LoRA |
|---|---|---|---|
| Format, 2 | |||
| OpenBookQA, 2 | |||
| OpenBookQA, 16 | |||
| Fixed output basis, rank 2 | OpenBookQA / Qwen-1.5B | Format / Llama-1B |
|---|---|---|
| Top spectral band | 46.52 | 64.43 |
| Bottom spectral band | 42.20 | 59.22 |
| Random orthonormal | 45.92 | 65.34 |
| Initial gradient, 64 sequences | 42.72 | 66.58 |
| Initial gradient, all sequences | 43.04 | 66.22 |
| After 10% warm-up | 51.16 | 72.21 |
| Task | Warm-up fraction | Direction 0 | Direction 301 | Direction 302 |
|---|---|---|---|---|
| OpenBookQA | 5% | 0.13 | 0.00 | |
| 10% | 0.58 | |||
| 25% | 0.78 | 0.70 | 0.81 | |
| Format | 5% | 0.40 | 0.45 | 0.23 |
| 10% | 0.63 | |||
| 25% | 0.93 | 0.90 | 1.07 |
| Question | Evidence used here | Inferential status |
|---|---|---|
| Is the rank-2 factor contrast larger than the top/bottom contrast? | Four complete A/B comparisons | Post hoc descriptive comparison; seed-bootstrap intervals for component contrasts |
| Can count or module allocation explain the advantage? | Count reversal; within-module controls | Two-band families for count reversal; eight-test family for module controls |
| Does the result survive selected training controls? | Fresh seeds, one GPU model, four-times tokens, answer-only loss | Prespecified campaign comparisons; scope differs by control |
| Does the A-frozen advantage need spectral placement of ? | Random orthonormal A, two pairs, rank 2 | Prespecified replication; both P1 tests pass two-pair Bonferroni; earlier outcome partly known |
| Does the contrast weaken with rank? | Complete rank sweep on one pair; rank-16 checks | Descriptive trend; no universal crossover or strict monotonicity claim |
| Does the rank-16 deficit persist with PEFT MiCA? | Transferred recipe and shared-recipe follow-up | Nominal paired tests; shared-recipe test planned after the first result |
| Historical comparison | Seeds | Observed difference | MDE |
|---|---|---|---|
| Format / Llama, minus top | 7 | 2.21 | 3.36 |
| Format / Llama, minus | 7 | 0.98 | 4.51 |
| Format / Llama, minus random orthonormal | 7 | 2.42 | 3.92 |
| Same random comparison, fresh seeds | 8 | 3.36 | 4.84 |
| Format / Llama, common-rate bottom advantage, | 5 | 0.98 | 3.29 |
| Format / Llama, common-rate bottom advantage, | 5 | 2.25 | 3.50 |
| Follow-up | Pushed (2026) | Scope and departures |
|---|---|---|
| Boundary selection | 03 Oct, 00:51 | Reuses prior selection runs; three engine fixes precede all new runs. |
| Single GPU model | 03 Oct, 00:51 | Seven arms rerun at their reported rates, seeds 2–6. |
| PEFT MiCA | 03 Oct, 00:51 | Before runs, addendum distinguishes recipes and training amounts. |
| Shared MiCA recipe | 03 Oct, 08:36 | Planned after the MiCA result; five new LoRA runs reuse MiCA controls. |
| Module groups | 03 Oct, 00:51 | Validation uses 500 rather than 400 items; some boundary optima and ties remain. |
| Warm-up directions | 03 Oct, 00:51 | New direction seeds 301/302; historical top controls and fixed reference gaps. |