LoRA fine-tuning adapts small language models (SLMs) to heterogeneous instruction data within a low-rank update subspace, making it vulnerable to three structural problems: conflicting gradients that cancel, static data selection that cannot track evolving learning dynamics, and subspace saturation that causes later updates to overwrite useful directions. We argue that effective adaptation therefore requires controlling which data-induced gradients enter the LoRA subspace and when. We propose GRADE (GRadient-Aligned Data-centric rEcipe), a data-centric framework combining two mechanisms: a state-aware selector that continually admits samples aligned with the evolving multi-task gradient field, and a self-calibrating step-level gate that rejects updates likely to cause destructive overwrite near saturation. Across three current-generation backbones and a heterogeneous seven-dataset instruction pool, GRADE outperforms strong data-selection and PEFT-stabilization baselines in accuracy and robustness. It is the only method to improve consistently over standard LoRA on every architecture, while producing more coherent gradient trajectories and less destructive overwrite. These results show that successful SLM adaptation depends not only on which data are selected, but also on which gradients are allowed to enter and persist in the constrained update subspace.
Figures & tables
Figure 1: Per-sample LoRA gradient cosine on Gemma-2-9B: training-pool datasets have inconsistent gradient directions (intra-dataset +0.029 , inter-dataset −0.003 ). Setup in App. O .
Figure 2: Overview of GRADE. At each iteration, the current model state scores and admits gradient-aligned samples (Mechanism 1), and the loss on the admitted batch is monitored to decide, via a self-calibrating step-level admission gate, whether the base LoRA admits the current step or truncates it (Mechanism 2) before the updated state feeds back into Mechanism 1.
Llama-3.1-8B (3t)
Qwen3-8B (3t)
Gemma-2-9B (3t)
Cross-arch
Method
Avg ±std
Δ
Avg ±std
Δ
Avg ±std
Δ
Worst Δ
Mean Δ
LoRA-MGPO
.6607±.0031
+0.85
.6543±.0048
− 0.28
.6948±.0027
+1.84
− 0.28
+0.80
Sensitivity-LoRA
.6551±.0042
+0.29
.6542±.0039
− 0.29
.6778±.0031
+0.14
− 0.29
+0.05
GRAD-MATCH
.6541±.0038
+0.19
.6543±.0041
− 0.28
.6837±.0029
+0.73
− 0.28
+0.21
PCGrad
.6568±.0035
+0.46
.6559±.0038
− 0.12
.6897±.0026
+1.33
− 0.12
+0.56
CAGrad
.6579±.0032
+0.57
.6564±.0034
− 0.07
.6912±.0023
+1.48
− 0.07
+0.66
Table 1: Main results on the heterogeneous pool ( n≈21K ): eight methods on three base backbones, ranked on the 3-task held-out average (ARC-C + HellaSwag + Winogrande). Δ in pp vs. LoRA; Worst- Δ is the minimum across architectures, Mean- Δ the unweighted mean. Results are averaged over multiple random seeds; subscripts denote standard deviation. Per-architecture and cross-architecture leaders bolded. Per-backbone learning rates and the ClusterUCB † reproduction note are in App. T .
Figure 3: Selection mechanism. (a) GFCintra/GFCinter coherence ratio across epochs (log- y , LLaMA-2-7B mechanism diagnostic): selection-equipped runs ( S1 adaptive, S2 frozen) reach 46 – 52× , while the no-selection LoRA baseline ( S3 ) collapses to 6.9× as GFCinter rises ∼4× on the full heterogeneous pool. (b) Per-dataset selection rate under loss-keyed (red) vs. alignment-keyed (blue) top-half cuts on the noisy heterogeneous pool. (c) Per-dataset selection-rate trajectory across epochs on Gemma-2-9B under alignment-keyed selection; dolly15k and xsum highlighted as the anti-correlated pair, the rest dimmed: selection is non-stationary and tracks the evolving model state.
Figure 4: Evidence for the GRADE admission gate. (a) Step-level dynamics on Llama-3.1-8B: 100-step-smoothed probe-loss EMA (blue, left axis) and cumulative skip fraction (red, right axis); gray dotted line marks the 0.5 stationary-noise prediction. (b) Cross-architecture gate ablation: 3-task average accuracy for gate-OFF (gray) vs gate-ON (blue) on the three backbones. (c) Per-task signed Δ heatmap. Gate-ON wins the 3-task average on all three backbones.
Figure 5: Admission Necessity Matrix: Δ vs (on, on) GRADE (pp) across the four corners of the { sample, step } admission lattice at matched lr =2 e − 5.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Backbone
cosintra ( t=0 )
cosinter ( t=0 )
r⋆raw
Llama-3.1-8B
0.2042
0.0097
0.779
Qwen3-8B
0.2114
0.0579
0.378
Gemma-2-9B
0.1392
0.0316
0.423
Appendix
Table 2: Measured (cosintra,cosinter) at t=0 on the heterogeneous pool ( T=7 datasets, auto_frac_n_samples =20 ), and the resulting closed-form r⋆ . All three rows produce cosinter>0 , so the safeguards in _r_star are inactive; the floor / hard-fallback / clamp serve as numerical guards rather than as load-bearing components of the recipe.
Figure 6: Operating regime of Eq. ( 2 ) on (cosintra,cosinter) . Colored markers: Apr-2026 main-table backbones (Tab. 2 ); gray markers: alternate-LR / smoke-test runs. Diagonal dotted line is the aligned-dataset limit cosinter=cosintra (where r⋆→1/T ); horizontal dashed line is the orthogonal-dataset limit cosinter=0 (where r⋆→1 ); the shaded cosinter≤0 band is where the floor would activate. Light gray contours are iso- r⋆ curves at r⋆∈{0.3,0.5,0.8} . All eleven measurements lie in the well-posed positive- cosinter region, so the safeguards in _r_star are never load-bearing in the reported numbers.
Figure 7: Selection–state coupled dynamics (LLaMA-2-7B mechanism diagnostic). (a) Admitted-population retention (Jaccard) across consecutive epochs for adaptive (top) vs. frozen (bottom) selection. (b) Selection-mechanism dynamics on a normalized axis: criterion sharpness (admitted/rejected score gap, blue) and population churn (Jaccard, red), both normalized to their epoch 0 value.
Figure 8: Routing-statistic monotonicity vs. noise on the four instrumented Phase-1.5 runs. The desirable region is the lower-right corner (monotone w.r.t. step, quiet within an epoch); Lprobe(t) is the unique statistic in that region. ρ^(t) and its log variant occupy the upper-left.
Figure 9: Per-dataset admission-gate gain vs. residual ∥gLoRA∥ (left) and the same scatter with ARC-C excluded (right). The OLS-through-origin slope drops from α≈0.0022 to α≈0.0012 but stays positive; Dolly remains the high-norm / no-trigger diagnostic deviation in both panels.
Figure 10: Probe vs. exact per-sample LoRA gradient, end-to-end. (a) Cost–loss Pareto: held-out loss vs. wall-clock cost for the forward probe and the exact vmapped gradient (mean ±σ , 3 seeds). (b) Seed-paired held-out loss for the two estimators across 3 seeds.
Figure 11: Per-sample alignment vs. loss on the Phase-1 noisy heterogeneous pool (Llama-3.1-8B, end of epoch 1 , n=19,119 ). Anchors gsm8k + arc in blue, the 5 instruction-style datasets in gray; the dashed lines are the loss-median (red, vertical) and alignment-median (blue, horizontal). The geometric companion to Figure 3 (b) of the main text.
Figure 12: Pool-fit improvement vs. the no-selection reference for three selection rules on the noisy heterogeneous pool (left) and the matched clean pool (right). Loss-based selection is the only rule whose noisy-pool improvement is negative; the gap collapses on the clean pool, identifying the failure as a noise-specific pathology. The per-sample geometric signature of this pathology is in Figure 3 (b).
Figure 13: Per-epoch gap, on a log- y axis, between the mean alignment score of selected vs. rejected samples (blue, probe units) and the absolute mean CE-loss gap of the same two populations (red, nats); dashed line at y=1 separates the two decades. Result analysis appears in App. S .
Llama-3.1-8B
Qwen3-8B
Gemma-2-9B
GRADE configuration
Avg
Δ
Avg
Δ
Avg
Δ
GRADE, lr =2×10−4
0.6568
+0.46
0.6589
+0.18
0.6828
+0.64
GRADE, lr =2×10−5
0.6602
+0.80
0.6550
− 0.21
0.6964
+2.00
Appendix
Table 3: GRADE under two learning-rate operating points on each backbone of Table 1 . The main-table GRADE row uses, per column, the configuration with the higher Avg — lr =2×10−5 on Llama-3.1-8B and Gemma-2-9B; lr =2×10−4 on Qwen3-8B. The Qwen3-8B and Gemma-2-9B preferences point in opposite LR directions, which is why a single shared learning rate would underperform a per-backbone choice on at least one backbone. Per-column Avg leaders are bolded.
Recent literature on fine-tuning Large Language Models highlights a fundamental debate. While Full Fine-Tuning (FFT) provides greater representational plasticity, Low-Rank Adaptation (LoRA) can match or surpass FFT performance while constraining updates to a low-rank space and potentially benefiting from additional regularization. Through empirical evaluation across diverse tasks (SQL, Medical QA, and Counterfactual Knowledge) and varying language models (Gemma-3-1B, Qwen2.5-1.5B, and Qwen2.5-3B), we observe both trends and find that the better static architecture depends on the task and model. Spectral and truncation analyses further show that endpoint compressibility alone does not explain these task differences, suggesting task-score sensitivity and constrained optimization trajectories as possible explanations. To address this challenge, we propose a Mixture of LoRA and Full (MoLF) Fine-Tuning, a unified framework that enables continuous navigation between both training regimes. MoLF dynamically routes updates between FFT and LoRA at the optimizer level to ensure that exact gradient signals are available to both experts throughout training, while only selected experts update their weights. For memory-constrained environments, we also introduce MoLF-Efficient, which freezes base weights and only routes updates among a pair of LoRA experts of potentially varying rank. Our evaluations show that MoLF either improves on or stays within 1.5 percentage points of the better of FFT and LoRA across the nine tested settings, while MoLF-Efficient outperforms both AdaLoRA and AdaMix in eight of nine settings, with gains over the stronger baseline of up to 11.70 percentage points on Fact, 3.13 on Med, and 2.98 on SQL.
Haozhan Tang, Xiuqi Zhu, Xinyin Zhang +3
Carnegie Mellon University · Tsinghua University · Infinigence AI
Data attribution methods using gradient similarity are widely used to analyze and select training data for large language models, but what gradient similarity actually measures is debated. Some interpret it as identifying task-relevant skills, while other work reports that surface form is the main factor. We resolve this debate for supervised fine-tuning examples by varying task and answer format independently. Specifically, we render benchmarks in different answer formats, such that datasets can share a task without a format or a format without a task. We find that gradient alignment follows the answer format, as benchmark pairs sharing an answer format align strongly (disattenuated cosine near 0.4), while same benchmarks rendered with different answer format classes show no alignment (near 0.0). We demonstrate that this ordering holds from the earliest pretraining checkpoints through post-training, and across model scales and families. We then analyze the released selections of LESS, a gradient-based data selection method for instruction tuning, and find that each target's selections over-represent the target's own answer format. Hence, we demonstrate that gradient-based attribution methods track format similarity more than task semantics, meaning that such methods, as well as the semantic interpretation of the gradient, should be tested on data where answer format and task vary independently for greater robustness and reliability.
Online batch selection fine-tunes a language model on the most useful part of each candidate batch. Selectors that match the gradient of the candidate batch are attractive because they need no held-out data, yet they rarely beat training on the whole batch. We show why. In-sample gradient matching uses every example as part of its own target, so its objective credits each example with its own gradient noise. This is the covariance penalty that makes training error optimistic, now sitting on the diagonal of the gradient Gram matrix: it steers selection toward the noisiest examples and makes the full batch the best solution the objective can reach. The fix costs nothing. For each example, the other candidates form an independent sample of the data distribution, so removing the diagonal turns the matching objective into an unbiased estimate of the update's error with respect to the population gradient. The minimizer of this leave-one-out objective weights examples by their gradient signal-to-noise ratio (SNR), and whenever per-example SNR is heterogeneous enough, half of a batch yields a lower-error update than the whole batch; we give the exact condition. We build \method{} on this principle. It computes the Gram matrix in the metric of the Adam preconditioner during the ordinary backward pass, selects a weighted subset greedily with a (1−e−γ) guarantee, and uses no held-out data. Across four fine-tuning tasks and seven backbones from 1.5B to 8B parameters, LOOM improves on full-batch training by 2.3 and 2.4 points on Llama-3.1-8B and Qwen2.5-7B, exceeds every in-sample gradient matcher by 2.4 points and the validation-guided GREATS and OPUS by 1.6--2.0, and selects injected label noise at under a fifth of its base rate.
Hongyu Chen, Xinyi Luo, Ming Zhao +6
Sichuan University, Chengdu, China · University of Electronic Science and Technology of China, Chengdu, China