Looped language models separate computational depth from parameter count by repeatedly applying the same transformer block. Adapting these models requires a shared update that remains effective as hidden states evolve throughout the recurrent computation. Our empirical analysis reveals a pronounced late-loop bias in standard low-rank adaptation (LoRA): the shared update provides limited adaptation at early loop positions. This imbalance motivates training shared updates under varying combinations of their applications. Randomly omitting adapter applications alone, however, does not improve task performance; it reduces expected update strength during training while leaving inference unchanged. We introduce Loop Dropout, which couples stochastic masking of adapter applications with inverse-survival rescaling to preserve expected update strength and promote effective adaptation across loops. Extensive experiments demonstrate improved mathematical reasoning across model sizes, adapter ranks and training recipes, with benefits extending to general instruction tuning and code generation. Loop Dropout outperforms existing LoRA variants and adapter regularizers, while further analysis shows stronger early-loop adaptation. Every backbone loop remains active, and inference applies the adapter at all loops using standard LoRA without additional trainable parameters or inference computation. Code is available at https://github.com/NUS-HPC-AI-Lab/loop-dropout .
Figures & tables
Figure 1: Loop Dropout reduces the late-loop bias of shared adaptation. Ouro ( Zhu et al., 2025 ) uses four loops through a shared transformer block by default. To probe adaptation across loops after GSM8K fine-tuning, we activate the adapter in one loop at a time while retaining all four backbone loops. LoRA’s loss reduction falls from 32% at the fourth loop to 12% at the first; Loop Dropout maintains 31–36% across positions. All reductions are relative to the frozen model and use final-loop teacher-forced test loss.
Figure 2: Loop Dropout regularizes shared adaptation across loops. The shared block F reuses the same W and Δ over K loops. Following Eq. ( 3 ), training uses gt=bt/q , where q=1−p and bt∼Bernoulli(q) independently for each example and loop. The training row illustrates one sampled pattern: retained updates have gate 1/q , and dropped updates have gate 0 . LoRA and inference use gt=1 at every loop. A zero gate removes only Δ ; every backbone loop remains active.
Ouro-1.4B
Ouro-2.6B
Benchmark
LoRA
LoRA+
Loop Dropout
LoRA
LoRA+
Loop Dropout
GSM8K
85.34 ± 0.64
85.54 ± 0.23
86.53 ± 0.44
87.62 ± 0.52
87.21 ± 0.38
88.73 ± 0.31
MATH-500
38.87 ± 1.67
38.13 ± 0.42
47.53 ± 2.61
42.87 ± 1.17
43.40 ± 0.53
47.20 ± 0.53
Table 1: Mathematical reasoning performance. Accuracy in %, mean ± SD over three training seeds per method. Bold marks the highest mean within each model and benchmark.
Method
GSM8K
MATH-500
Train (min)
Memory (GB)
LoRA
85.34 ± 0.64
38.87 ± 1.67
110.6
45.36
LoRA+
85.54 ± 0.23
38.13 ± 0.42
112.8
45.40
CoTo-on-Ouro
85.60 ± 0.76
37.80 ± 1.06
94.1
45.28
LoRA Dropout
85.77 ± 0.70
42.13 ± 0.31
444.9
45.37
Loop Dropout
86.53 ± 0.44
47.53 ± 2.61
128.9
45.36
Table 3: Comparison with existing adaptation methods. Ouro-1.4B, rank 16, mathematical recipe. Accuracy is mean ± SD over three training seeds; training time and peak memory are their means. Displayed costs use an H100 80GB and exclude learning-rate search and evaluation. All methods use 15.1M adapter parameters.
Method
Only loop 1
Only loop 2
Only loop 3
Only loop 4
LoRA
0.03 ± 0.04
0.30 ± 0.46
2.63 ± 3.07
6.07 ± 4.42
Loop Dropout
1.31 ± 0.69
64.72 ± 13.56
86.23 ± 0.84
86.18 ± 0.78
Table 4: GSM8K generation with one adapter application. Ouro-1.4B after MetaMath-GSM fine-tuning; accuracy in %, mean ± SD over three seeds. Only the indicated adapter application is enabled, with all four backbone loops retained.
Method
GSM8K
MATH-500
LoRA
85.32 ± 0.92
40.13 ± 1.62
Unscaled
84.71 ± 0.19
39.33 ± 1.70
Dose control
85.57 ± 0.88
38.60 ± 1.71
Module-wise
85.52 ± 0.62
41.13 ± 0.76
Low-rank weight noise
85.70 ± 0.54
39.13 ± 0.83
Parallel noise
85.62 ± 0.64
37.93 ± 1.33
Table 5: Ablating loop-level masking and inverse-survival rescaling. The controls separate update scaling, gate variation across loops, masking granularity and matched noise. All rows use shared rank-16 adapters on Ouro-1.4B, retain four backbone loops and apply the full adapter update at inference. MetaMath-GSM recipe, learning rate 10−4 ; accuracy in %, mean ± SD over three seeds. Noise strength c=1 matches the expected perturbation energy of Loop Dropout. Per-seed results are in Tables 10 and 15 .
Without loop mask
With loop mask
Adapters
Rank
Params
GSM8K
MATH-500
GSM8K
MATH-500
Shared
16
15.1M
84.69
37.00
87.04
48.40
Independent
4
15.1M
84.08
39.60
87.57
44.40
Independent
16
60.6M
84.53
39.40
86.43
43.80
Table 6: Loop Dropout with shared and independent adapters. Shared adapters reuse their parameters across all four loops; Independent adapters learn separate parameters for each loop. Rank is per loop, and Params counts all trainable adapter parameters. With loop mask uses complete-application masking and inverse-survival rescaling during training; all adapters are active at inference. MetaMath-GSM recipe with development selection over five learning rates per method; accuracy in % for training seed 101.
Figure 3: Mathematical accuracy across adapter ranks. Ouro-1.4B after MetaMath-GSM fine-tuning. Markers and horizontal bars show mean ± SD over three seeds. Both methods use the same five-rate search budget at each rank, with α/r=2 ; Table 20 gives the values.
Figure 4: Depth transfer on Ouro. Loop Dropout gains over LoRA on GSM8K (percentage points). Outlined cells match training and evaluation depth.
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
MetaMath-GSM-100k
GSM8K
Tülu 2-100k
Training examples
100,000
6,973
98,415
Epochs / optimizer steps
1 / 3,125
2 / ≈ 1,744
1 / 769
Effective batch
8×4=32
8×1=8
8×16=128
Maximum sequence length
1,024
512
2,048
Warm-up
3%
5%
3%
Learning rates
See below
3e-5 to 6e-4
1e-4
Appendix
Table 7: Fine-tuning recipes. Effective batch = micro-batch × gradient accumulation. The adapter recipes use AdamW ( Loshchilov and Hutter, 2019 ) ( β1=0.9 , β2=0.999 ), zero weight decay, gradient clipping at 1.0, a cosine schedule decaying to 10% of the peak learning rate, and the final checkpoint.
Method
Per-seed time
Mean ± SD
Memory (GB)
LoRA
111.93 / 106.74 / 113.28
110.65 ± 3.45
45.36
LoRA+
111.20 / 107.74 / 119.40
112.78 ± 5.99
45.40
CoTo-on-Ouro
94.95 / 93.32 / 94.06
94.11 ± 0.82
45.28
LoRA Dropout
442.39 / 438.34 / 453.98
444.91 ± 8.12
45.37
Loop Dropout
130.86 / 128.16 / 127.65
128.89 ± 1.72
45.36
Appendix
Table 8: Training time across seeds on H100 80GB. Times are in minutes for seeds 101 / 102 / 103, with their mean and sample SD. Peak memory is the mean over the same runs.
LoopUS-Qwen3-4B
Huginn
Benchmark
LoRA
Loop Dropout
LoRA
Loop Dropout
GSM8K
83.93
85.52
59.89
60.11
MATH-500
34.60
40.80
13.00
13.80
Appendix
Table 9: Mathematical evaluation on other looped transformers. Accuracy in %. Bold marks the higher value within each model and benchmark.
Method
LR
GSM8K seeds 101 / 102 / 103
MATH-500 seeds 101 / 102 / 103
LoRA
1e-4
84.6 / 85.0 / 86.4
39.2 / 42.0 / 39.2
Loop Dropout
1e-4
87.7 / 87.0 / 88.0
44.0 / 45.4 / 42.6
Unscaled
1e-4
84.5 / 84.7 / 84.9
40.6 / 40.0 / 37.4
Unscaled
2e-4
83.8 / 83.8 / 84.8
35.6 / 38.0 / 37.4
Dose control
1e-4
85.4 / 84.8 / 86.5
36.8 / 40.2 / 38.8
Dose control
2e-4
85.5 / 84.9 / 85.9
37.8 / 42.4 / 37.8
Appendix
Table 10: Per-seed results of the mask-control comparison on Ouro-1.4B under the MetaMath protocol, at the learning rates shown; accuracy in %.
Metric
LoRA
Loop Dropout
MMLU (0-shot)
68.64 ± 0.40
68.91 ± 0.13
BBH (3-shot CoT)
71.20 ± 0.14
70.78 ± 0.11
TruthfulQA MC2
47.32 ± 1.02
48.19 ± 0.68
HumanEval
71.54 ± 2.14
72.56 ± 0.61
HumanEval+
67.68 ± 3.05
69.92 ± 0.35
MBPP
75.57 ± 0.40
76.72 ± 0.53
Appendix
Table 11: All downstream metrics after Tülu 2 instruction tuning of Ouro-1.4B. Rank 16; accuracy in %, mean ± SD over three matched seeds. HumanEval+ and MBPP+ add the expanded EvalPlus tests to the original ones; IFEval reports prompt-level accuracy. The last three rows average the four code metrics, the six-benchmark set, and all nine metrics per seed; bold marks the highest average in each summary row.
Variant
Mask
LR 1e-4
LR 3e-4
LoRA
none
80.6 ± 0.3
78.1 ± 1.0
Loop Dropout
uniform p=0.5
81.7 ± 0.2
81.3 ± 0.7
Early-heavy profile
p=(.75,.75,.25,.25)
82.1 ± 0.2
79.4 ± 2.6
Late-heavy profile
p=(.25,.25,.75,.75)
80.0 ± 0.5
77.9 ± 1.0
Ramp down
p=(.8,.6,.4,.2)
80.7 ± 0.8
79.3 ± 0.4
Ramp up
p=(.2,.4,.6,.8)
80.0 ± 0.8
79.0 ± 2.6
Appendix
Table 12: Structured, scheduled and adaptive masks on the GSM8K recipe (Ouro-1.4B, strict exact match, %; two seeds, mean ± SD). Masks control adapter applications while all four backbone loops run; all variants keep the 1/P(keep) rescaling. Bottom: five-seed follow-up of the two closest variants. No variant exceeds the uniform rule at both learning rates.
Rank
Benchmark
LoRA
Loop Dropout
Difference
4
GSM8K
85.82 ± 0.20
86.38 ± 0.50
+0.56
4
MATH-500
39.00 ± 1.25
46.07 ± 1.17
+7.07
64
GSM8K
85.44 ± 1.12
87.01 ± 0.49
+1.57
64
MATH-500
37.87 ± 1.70
43.33 ± 1.85
+5.47
128
GSM8K
85.06 ± 0.62
86.48 ± 0.22
+1.42
128
MATH-500
34.80 ± 1.00
41.60 ± 1.39
+6.80
Appendix
Table 13: Paired comparisons at different shared adapter ranks. Three seeds, LR 10−4 and α/r=2 . Accuracy is in %; differences are in percentage points.
Missing answers (%)
Net gain (points)
Rank
LoRA
Loop Dropout
Both extractable
Missing in either
Total
4
8.80
9.60
6.80
0.27
7.07
64
7.60
5.60
4.80
0.67
5.47
128
7.80
5.67
6.00
0.80
6.80
Appendix
Table 14: Decomposing the net MATH-500 accuracy difference. Missing-answer rates are percentages. Both net-difference components use all 500 problems as denominator and sum to the overall gain in percentage points. This decomposition does not change the official scorer.
Method
Rank
LR
c
GSM8K
MATH-500
LoRA
4
1e-4
–
85.90 / 85.60 / 85.97
38.0 / 38.6 / 40.4
Loop Dropout
4
1e-4
–
86.13 / 86.05 / 86.96
45.6 / 47.4 / 45.2
LoRA
64
1e-4
–
84.15 / 86.20 / 85.97
36.2 / 39.6 / 37.8
Loop Dropout
64
1e-4
–
87.57 / 86.66 / 86.81
44.4 / 41.2 / 44.4
LoRA
128
1e-4
–
84.91 / 84.53 / 85.75
35.8 / 33.8 / 34.8
Loop Dropout
128
1e-4
–
86.35 / 86.35 / 86.73
43.2 / 40.8 / 40.8
Appendix
Table 15: Per-seed results of the additional suites. Seeds are ordered 101 / 102 / 103. Accuracy is in %; c is the noise strength and is inapplicable to LoRA and Loop Dropout.
Seed
GSM8K
MATH-500
101
84.84
37.40
102
85.60
39.00
103
86.35
37.00
Mean ± SD
85.60 ± 0.76
37.80 ± 1.06
Appendix
Table 16: Per-seed CoTo-on-Ouro results. Full GSM8K and MATH-500 test sets under the MetaMath protocol; accuracy in %. The summary uses the sample standard deviation.
LoRA+
LoRA Dropout
Seed
GSM8K
MATH-500
GSM8K
MATH-500
101
85.29
37.80
85.44
42.20
102
85.60
38.60
85.29
41.80
103
85.75
38.00
86.58
42.40
Mean ± SD
85.54 ± 0.23
38.13 ± 0.42
85.77 ± 0.70
42.13 ± 0.31
Appendix
Table 17: LoRA+ and LoRA Dropout across three training seeds. Ouro-1.4B, rank 16, MetaMath-GSM recipe; accuracy in %.
Method
Params
GSM8K
MATH-500
Shared LoRA, rank 16
30.3M
87.62 ± 0.52
42.87 ± 1.17
LoRA+, rank 16
30.3M
87.21 ± 0.38
43.40 ± 0.53
Independent, rank 4
30.3M
85.95 ± 2.59
43.60 ± 0.40
Independent, rank 16
121.1M
85.87 ± 0.87
41.40 ± 1.71
Loop Dropout, rank 16
30.3M
88.73 ± 0.31
47.20 ± 0.53
Appendix
Table 18: Shared and independent adapters on Ouro-2.6B. MetaMath-GSM recipe, mean ± SD over three training seeds, in %. All backbone loops and all trained adapter applications are active at inference.
Method
LR
GSM8K
MATH-500
Shared LoRA, rank 16
2×10−4
88.02 / 87.04 / 87.79
42.00 / 44.20 / 42.40
LoRA+, rank 16
10−4
87.04 / 86.95 / 87.64
43.80 / 42.80 / 43.60
Independent, rank 4
2×10−4
86.88 / 87.95 / 83.02
43.20 / 44.00 / 43.60
Independent, rank 16
2×10−4
85.90 / 84.99 / 86.73
39.80 / 41.20 / 43.20
Loop Dropout, rank 16
10−4
89.01 / 88.78 / 88.40
47.80 / 47.00 / 46.80
Appendix
Table 19: Per-seed Ouro-2.6B results after independent tuning. Seeds are ordered 101 / 102 / 103; accuracy in %.
Rank
Method
GSM8K
MATH-500
4
LoRA
85.82 ± 0.20
39.00 ± 1.25
4
Loop Dropout
86.38 ± 0.50
46.07 ± 1.17
16
LoRA
85.34 ± 0.64
38.87 ± 1.67
16
Loop Dropout
86.53 ± 0.44
47.53 ± 2.61
64
LoRA
85.44 ± 1.12
37.87 ± 1.70
64
Loop Dropout
85.87 ± 0.48
40.73 ± 1.42
Appendix
Table 20: Independent tuning at each adapter rank. Ouro-1.4B, three seeds per rank and method; full GSM8K and MATH-500 evaluation, mean ± SD in %. Both methods have five learning-rate candidates at each rank.
Rank
Method
LR
GSM8K
MATH-500
4
LoRA
10−4
85.90 / 85.60 / 85.97
38.00 / 38.60 / 40.40
4
Loop Dropout
10−4
86.13 / 86.05 / 86.96
45.60 / 47.40 / 45.20
16
LoRA
2×10−4
84.69 / 85.37 / 85.97
37.00 / 40.20 / 39.40
16
Loop Dropout
5×10−5
87.04 / 86.35 / 86.20
48.40 / 49.60 / 44.60
64
LoRA
10−4
84.15 / 86.20 / 85.97
36.20 / 39.60 / 37.80
64
Loop Dropout
2×10−4
85.60 / 85.60 / 86.43
42.00 / 39.20 / 41.00
Appendix
Table 21: Per-seed results after independent rank-wise tuning. Seeds are ordered 101 / 102 / 103; accuracy in %. The same learning rate is used for all three seeds in a row.
Group
n
LoRA
Loop Dropout
Difference
Algebra
124
58.33 ± 3.05
72.04 ± 3.05
+13.71
Counting & Probability
38
29.82 ± 9.24
39.47 ± 2.63
+9.65
Geometry
41
32.52 ± 6.14
38.21 ± 3.73
+5.69
Intermediate Algebra
97
19.59 ± 1.03
25.43 ± 6.21
+5.84
Number Theory
62
40.32 ± 5.59
44.09 ± 4.06
+3.76
Prealgebra
82
54.88 ± 4.40
62.60 ± 1.86
+7.72
Appendix
Table 22: MATH-500 breakdown for independently tuned rank-16 adapters. Mean ± SD over three seeds, in %; differences in percentage points. Each partition covers all 500 problems, with n denoting the number of problems per group. Checkpoints are identical to those in Table 20 .
Figure 5: Zero-shot generalization beyond the training depth. Ouro-1.4B after direct GSM8K fine-tuning at four loops, evaluated without further training; accuracy is mean ± SD over three seeds. Annotations show the gain of Loop Dropout over LoRA in percentage points.
Method
Train K
Eval K=4
Eval K=6
Eval K=8
LoRA
4
79.9 ± 1.0
75.0 ± 1.0
70.8 ± 1.9
LoRA
6
82.4 ± 1.1
82.3 ± 0.7
77.8 ± 0.7
LoRA
8
80.3 ± 1.4
81.8 ± 0.6
81.6 ± 0.9
Loop Dropout
4
81.8 ± 0.5
78.0 ± 1.7
75.6 ± 1.3
Loop Dropout
6
83.0 ± 0.3
83.3 ± 0.7
80.5 ± 0.8
Loop Dropout
8
83.4 ± 0.5
83.6 ± 0.5
82.1 ± 0.8
Appendix
Table 23: Accuracy across training and evaluation depths. GSM8K recipe, Ouro-1.4B, LR 10−4 ( A -factor rate for LoRA+), seeds 101–103; mean ± SD of strict exact match, %. Loop Dropout has higher mean accuracy than LoRA in all nine depth pairs. Comparisons within a training depth share the training budget; different training depths use different numbers of recurrent passes per optimizer step.
Model
Method
LR
GSM8K
Ouro-1.4B
LoRA
1e-4
79.9 ± 1.0
Ouro-1.4B
LoRA+
1e-4
79.8 ± 0.5
Ouro-1.4B
rsLoRA
1e-4
79.8 ± 0.9
Ouro-1.4B
Loop Dropout
1e-4
81.8 ± 0.5
Ouro-1.4B
LoRA
1e-4
80.4 ± 0.6
Ouro-1.4B
Loop Dropout
1e-4
81.9 ± 0.4
Appendix
Table 24: The gain persists across learning rates and model sizes under the GSM8K recipe. Zero-shot strict exact match (%) after fine-tuning on the GSM8K training split, mean ± SD. The first and last blocks use three seeds; the middle block uses seeds 0–4. LR denotes the A -factor rate for LoRA+. Bold marks the highest mean within each model, seed set and learning rate.
LR
LoRA
Loop Dropout
Δ
3e-5
80.4 ± 0.4
80.9 ± 0.3
+0.56
1e-4
80.4 ± 0.6
81.9 ± 0.4
+1.52
2e-4
79.8 ± 1.5
82.2 ± 0.6
+2.38
3e-4
78.6 ± 1.1
81.2 ± 0.6
+2.62
6e-4
72.4 ± 0.7
73.8 ± 0.8
+1.36
Appendix
Table 25: Learning-rate curve on the GSM8K recipe (Ouro-1.4B, K=4 ; lm-eval zero-shot strict exact match, %). Five seeds at 10−4 and 3×10−4 , three seeds elsewhere.
GSM8K-Platinum
GSM-Plus mini
Method
LR
Accuracy
Δ vs. LoRA
Accuracy
Δ vs. LoRA
LoRA
1e-4
82.4 ± 0.0
–
58.9 ± 0.7
–
Input dropout
1e-4
81.7 ± 0.4
−0.66
59.1 ± 0.4
+0.21
Loop Dropout
1e-4
83.8 ± 0.4
+1.38
60.0 ± 0.2
+1.14
LoRA
3e-4
80.3 ± 0.4
–
56.0 ± 1.0
–
Input dropout
3e-4
80.7 ± 1.7
+0.33
57.0 ± 1.1
+1.04
Appendix
Table 26: Transfer of the saved GSM8K-recipe adapters (Ouro-1.4B, three seeds) to GSM8K-Platinum (1,209 relabelled problems) and GSM-Plus mini (2,400 perturbed problems); zero-shot strict exact match, %. “Input dropout” applies standard dropout of 0.1 to the LoRA input.
K
Seeds
LoRA
Loop Dropout
Δ
2
0–2
68.6 ± 0.8
70.7 ± 0.3
+2.17
4
0–4
80.4 ± 0.6
81.9 ± 0.4
+1.52
6
0–2
81.6 ± 0.6
83.2 ± 1.2
+1.59
4
101–103
79.9 ± 1.0
81.8 ± 0.5
+1.90
6
101–103
82.3 ± 0.7
83.3 ± 0.7
+1.06
8
101–103
81.6 ± 0.9
82.1 ± 0.8
+0.51
Appendix
Table 27: Training and evaluating at the same recurrence depth on the GSM8K recipe (Ouro-1.4B, LR 10−4 ; strict exact match, %). Seeds 0–2 (five at K=4 ) use independent initializations; seeds 101–103 share initialization across methods.
Final-loop loss
All-on
Method
LR
Only 1
Only 2
Only 3
Only 4
accuracy
LoRA
10−4
0.7437
0.6834
0.6208
0.5709
80.59
Unscaled
10−4
0.5309
0.5082
0.4984
0.5037
77.03
Loop Dropout
10−4
0.5847
0.5475
0.5409
0.5432
81.73
LoRA
3×10−4
0.7493
0.6835
0.6159
0.5317
77.86
Unscaled
3×10−4
0.5152
0.4907
0.4824
0.4895
73.84
Appendix
Table 28: Single-loop activation of a trained update. Ouro-1.4B, direct GSM8K recipe. The four loss columns are final-loop teacher-forced cross entropy (lower is better). The last column separately reports standard all-on generation accuracy on all 1,319 test questions, in %.
Readout
LoRA
Loop Dropout
Loop 1
1.240 ± 0.041
1.047 ± 0.010
Loop 2
0.836 ± 0.006
0.765 ± 0.005
Loop 3
0.742 ± 0.009
0.717 ± 0.002
Loop 4
0.737 ± 0.010
0.700 ± 0.005
Appendix
Table 29: MATH-500 readouts along the fully adapted recurrence. Ouro-1.4B after MetaMath-GSM fine-tuning. Token-weighted cross entropy, mean ± SD over three training seeds; lower is better. Every adapter application and all four backbone loops are enabled.
Figure 6: Generation and early readouts after MetaMath-GSM fine-tuning. Ouro-1.4B, rank 16, four backbone loops. Left: GSM8K generation with only the indicated adapter application enabled. Right: MATH-500 teacher-forced cross entropy with all applications enabled, read out at each loop. Points and error bars show mean ± SD over three training seeds. Loop Dropout supports generation from individual applications at loops two through four and lowers MATH-500 loss at every readout, most strongly at the first loop.
Low-rank adaptation (LoRA) is a widely used parameter-efficient fine-tuning method that places trainable low-rank adapters into frozen pre-trained models. Recent studies show that using fewer LoRA adapters may still maintain or even improve performance, but existing methods still distribute adapters broadly, leaving where to place a limited number of adapters to maximize performance largely open. To investigate this, we introduce PAGE (Projected Adapter Gradient Energy), a gradient-based sensitivity probe that estimates the initial trainable gradient energy available to each candidate LoRA adapter. Surprisingly, we find that PAGE is highly concentrated on a single shallow FFN down-projection across two model families and four downstream tasks. We term this module the dominant adaptation module and show that its layer index is architecture-dependent but task-stable. Motivated by this finding, we propose DomLoRA, a placement method that places a single adapter at the dominant adaptation module. With only ~0.7% of vanilla LoRA's trainable parameters, DomLoRA outperforms it on average across various downstream tasks, including instruction following, mathematical reasoning, code generation, and multi-turn conversation. This method also improves other LoRA variants, supporting the dominant adaptation module perspective as a practical placement guideline.
Suoxin Zhang, Run He, Di Fang +3
South China University of Technology, China · Zhejiang University, China
Low-Rank Adaptation (LoRA) has become one of the most widely used fine-tuning mechanisms for adapting large language models to new domains, tasks, and users. Yet adaptation performance alone can obscure an important failure mode: LoRA updates may improve performance on the target distribution while degrading prior capabilities learned during pretraining and alignment. We show that this forgetting becomes especially severe when the adaptation distribution differs substantially from the models original training or alignment distributions. The challenge is amplified in practical settings, where the original training and alignment data are typically unavailable. Motivated by this constraint, we study how LoRA based adaptation balances new learning against forgetting in a replay-free setting, and introduce a simple output space regularizer that can be added directly to existing training pipelines. Our method removes the ground-truth token from both the base and adapted model distributions, renormalizes the remaining probabilities, and applies KL regularization only over the non-target vocabulary. This preserves the base models relative preferences among alternative tokens without directly opposing the cross-entropy signal required for adaptation. As the regularizer acts only at the loss level, it requires no replay data, architectural changes, adapter redesign, or inference-time overhead, and can be applied directly to existing LoRA variants. Across all LoRA variants tested and across various backbones, our method improves the frontier between new learning and forgetting when the adaptation distribution differs substantially from the base models original training or alignment distributions, suggesting a broadly applicable route toward more reliable LLM updating.
Runze Xu, Arpit Garg, Hemanth Saratchandran +1
Australian Institute for Machine Learning Adelaide University Adelaide SA 5000
Low-Rank Adaptation (LoRA) is the most widely adopted method for fine-tuning large language models. Notably, LoRA is inherently overparameterized: multiple pairs of low-rank factors can yield the same adapted weight matrix. We show--both theoretically and empirically--that these pairs exhibit significantly different condition numbers. As a result, converging to different loss minimizers directly impacts the convergence rate of LoRA. Building on this observation, we introduce Balanced Low-Rank Adaptation (BaLoRA), a variant of LoRA that projects iterates onto a balanced manifold. This manifold improves the conditioning of the loss landscape while preserving the adapted matrix. The projection step is computationally lightweight and integrates seamlessly into existing fine-tuning pipelines. Empirically, BaLoRA converges faster than standard LoRA and achieves superior performance across a range of fine-tuning tasks.
Valérie Castin, Kimia Nadjahi, Pierre Ablin +1
1 ´Ecole Normale Sup´erieure PSL, Paris, France · 2CNRS · 3Apple, Paris, France.