Every Batch Is Its Own Validation Set: Leave-One-Out Gradient Matching for Online Data Selection in LLM Fine-Tuning
Organizations: Sichuan University, Chengdu, China · University of Electronic Science and Technology of China, Chengdu, China
Abstract
Online batch selection fine-tunes a language model on the most useful part of each candidate batch. Selectors that match the gradient of the candidate batch are attractive because they need no held-out data, yet they rarely beat training on the whole batch. We show why. In-sample gradient matching uses every example as part of its own target, so its objective credits each example with its own gradient noise. This is the covariance penalty that makes training error optimistic, now sitting on the diagonal of the gradient Gram matrix: it steers selection toward the noisiest examples and makes the full batch the best solution the objective can reach. The fix costs nothing. For each example, the other candidates form an independent sample of the data distribution, so removing the diagonal turns the matching objective into an unbiased estimate of the update's error with respect to the population gradient. The minimizer of this leave-one-out objective weights examples by their gradient signal-to-noise ratio (SNR), and whenever per-example SNR is heterogeneous enough, half of a batch yields a lower-error update than the whole batch; we give the exact condition. We build \method{} on this principle. It computes the Gram matrix in the metric of the Adam preconditioner during the ordinary backward pass, selects a weighted subset greedily with a guarantee, and uses no held-out data. Across four fine-tuning tasks and seven backbones from 1.5B to 8B parameters, LOOM improves on full-batch training by 2.3 and 2.4 points on Llama-3.1-8B and Qwen2.5-7B, exceeds every in-sample gradient matcher by 2.4 points and the validation-guided GREATS and OPUS by 1.6--2.0, and selects injected label noise at under a fifth of its base rate.
Figures & tables
| Llama-3.1-8B | Qwen2.5-7B | ||||||||||
| Method | Ext. | MMLU | SciQA | GSM8K | HumanE. | Avg. | MMLU | SciQA | GSM8K | HumanE. | Avg. |
| Regular | – | 38.3 | 93.2 | 56.0 | 29.3 | 54.2 | 55.3 | 94.6 | 78.2 | 45.8 | 68.5 |
| Random | – | 35.6 | 92.9 | 54.9 | 26.8 | 52.6 | 54.6 | 93.5 | 77.8 | 41.3 | 66.8 |
| MaxLoss | – | 35.7 | 92.8 | 55.4 | 27.2 | 52.8 | 54.8 | 93.2 | 77.9 | 42.1 | 67.0 |
| MaxGrad | – | 35.9 | 92.8 | 55.1 | 26.9 | 52.7 | 54.7 | 93.9 | 77.7 | 41.6 | 67.0 |
| UDS | – | 40.1 | 94.3 | 58.9 | 30.8 | 56.0 | 59.6 | 95.2 | 79.6 | 46.2 | 70.2 |
| Variant | Avg. | Err. | |
| Loom (full) | 56.53 | — | 0.61 |
| Target | |||
| in-sample target (keeps diagonal) | 53.74 | 1.38 | |
| momentum target | 55.34 | 0.86 | |
| 2-fold cross-fitted target | 56.21 | 0.68 | |
| Metric | |||
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| Check | Result |
|---|---|
| Prop. 1 (ii): in-sample target error regressed on | slope (theory ) |
| Prop. 1 (i): mean leave-one-out target error ( ) | (theory ) |
| Prop. 1 (iii): risk estimate of a fixed separable rule, relative error | LOO ; in-sample (penalty ) |
| Prop. 2 : closed-form SNR weights vs. numerical optimum | max. relative difference |
| Prop. 2 : optimal risk vs. | vs. |
| Thm. 1 : oracle risk ratio at , MC vs. | vs. ; vs. ; vs. |
| Task | Train (size) | Test (size) |
|---|---|---|
| MMLU | auxiliary train (99,842) | test (14,042) |
| ScienceQA | train (12,726) | test (4,241) |
| GSM8K | train (7,473) | test (1,319) |
| Code | CodeAlpaca-20k (20,022) | HumanEval (164) |
| Hyperparameter | Value |
|---|---|
| LoRA rank / / dropout | 8 / 16 / 0 |
| Target modules | q, k, v, o, gate, up, down |
| Optimizer | AdamW ( , wd 0) |
| Learning rate | (MMLU), (others) |
| Schedule | warm-up ratio 0.01, cosine decay |
| Epochs | 1 (MMLU, GSM8K), 20 (ScienceQA), 2 (Code) |
| Llama-3.1-8B | Qwen2.5-7B | |||||||
|---|---|---|---|---|---|---|---|---|
| Method | MMLU | SciQA | GSM8K | HumanE. | MMLU | SciQA | GSM8K | HumanE. |
| Regular | ||||||||
| Random | ||||||||
| MaxLoss | ||||||||
| MaxGrad | ||||||||
| UDS | ||||||||
| Method | Q2.5-1.5B | L3.2-3B | Q2.5-3B | Mis-7B | Q3-8B |
|---|---|---|---|---|---|
| Regular | 57.6 | 45.9 | 63.3 | 51.8 | 75.9 |
| Random | 56.0 | 44.4 | 61.7 | 50.2 | 74.5 |
| GradMatch | 57.0 | 45.3 | 62.8 | 51.3 | 75.4 |
| UDS | 58.7 | 47.1 | 64.4 | 53.0 | 76.8 |
| GREATS | 58.1 | 46.4 | 63.8 | 52.3 | 76.3 |
| Loom | 59.5 | 47.8 | 65.2 | 53.7 | 77.6 |
| Backbone | Method | MMLU | SciQA | GSM8K | HumanE. | Avg. |
|---|---|---|---|---|---|---|
| Qwen2.5-1.5B | Regular | 44.8 | 88.1 | 61.9 | 35.4 | 57.6 |
| Random | 43.1 | 87.2 | 60.8 | 32.7 | 56.0 | |
| GradMatch | 44.3 | 87.8 | 61.5 | 34.4 | 57.0 | |
| UDS | 46.2 | 88.9 | 63.4 | 36.2 | 58.7 | |
| GREATS | 45.6 | 88.6 | 62.7 | 35.3 | 58.1 | |
| Loom | 47.1 | 89.4 | 63.9 | 37.4 | 59.5 |
| Method | 0% | 20% | 40% | Corr. sel. |
|---|---|---|---|---|
| Regular | 54.20 | 51.10 | 47.35 | (20.0) |
| Random | 52.55 | 49.60 | 45.70 | (20.0) |
| MaxLoss | 52.78 | 47.20 | 41.90 | 43.6 |
| UDS | 56.03 | 53.05 | 49.40 | 15.2 |
| CRAIG | 53.40 | 49.80 | 45.60 | 29.8 |
| GradMatch | 53.68 | 50.20 | 46.30 | 27.5 |
| Method | MMLU | SciQA | GSM8K | HumanE. | Avg. |
|---|---|---|---|---|---|
| Regular | 35.6 | 91.8 | 52.4 | 24.6 | 51.10 |
| Random | 33.4 | 91.1 | 51.1 | 22.8 | 49.60 |
| MaxLoss | 31.2 | 90.2 | 48.6 | 18.8 | 47.20 |
| UDS | 37.0 | 93.1 | 55.2 | 26.9 | 53.05 |
| CRAIG | 34.1 | 91.3 | 51.3 | 22.5 | 49.80 |
| GradMatch | 34.5 | 91.6 | 51.7 | 23.0 | 50.20 |
| Method | Thr. | Mem. | Ovh. | Steps | Time |
|---|---|---|---|---|---|
| Regular | 1.00 | 37.9 | – | 100% | 100% |
| Random | 1.74 | 31.6 | – | – | – |
| UDS | 1.21 | 32.8 | 2.9% | 62% | 51% |
| GradMatch | 0.95 | 41.2 | 4.8% | – | – |
| GREATS | 0.92 | 41.9 | 8.6% | 88% | 96% |
| Loom -fast | 1.17 | 33.4 | 3.6% | 64% | 55% |
| Stage | ms / step | Share |
|---|---|---|
| forward | 3,120 | 32.0% |
| backward | 6,030 | 61.9% |
| per-example gradients | 318 | 3.3% |
| Gram matrix | 142 | 1.5% |
| greedy selection + NNLS | 39 | 0.4% |
| optimizer update | 93 | 1.0% |
| Variant | Avg. | Err. | |
| Loom (full) | 56.53 | — | 0.61 |
| Target | |||
| in-sample target (keeps diagonal) | 53.74 | 1.38 | |
| momentum target | 55.34 | 0.86 | |
| 2-fold cross-fitted target | 56.21 | 0.68 | |
| Metric | |||
| Category | % |
|---|---|
| answer inconsistent with the question or options | 31 |
| truncated or malformed response | 19 |
| duplicate / near-duplicate of a batch neighbour | 14 |
| off-format (explanation without an answer letter) | 17 |
| correct and well-formed | 19 |
| Llama-3.1-8B | Mistral-7B-v0.3 | |||||||
|---|---|---|---|---|---|---|---|---|
| Method | MMLU | TydiQA | BBH | Avg. | MMLU | TydiQA | BBH | Avg. |
| Regular | 64.1 | 56.3 | 63.2 | 61.2 | 61.2 | 55.7 | 57.9 | 58.3 |
| Random | 63.4 | 55.1 | 62.0 | 60.2 | 60.3 | 54.6 | 56.8 | 57.2 |
| GradMatch | 63.8 | 55.8 | 62.7 | 60.8 | 60.8 | 55.2 | 57.4 | 57.8 |
| UDS | 64.9 | 57.2 | 63.9 | 62.0 | 61.9 | 56.5 | 58.6 | 59.0 |
| GREATS | 65.3 | 60.4 | 64.1 | 63.3 | 62.2 | 60.2 | 58.9 | 60.4 |
| Method | MMLU | SciQA | GSM8K | HumanE. | Avg. |
|---|---|---|---|---|---|
| ratio | 12.5% | 50% | 25% | 50% | |
| Regular | 55.3 | 94.6 | 78.2 | 45.8 | 68.5 |
| GradMatch | 53.9 | 94.3 | 77.4 | 44.5 | 67.5 |
| GREATS | 58.1 | 94.2 | 78.6 | 45.1 | 69.0 |
| UDS | 63.2 | 95.2 | 79.9 | 46.2 | 71.1 |
| Loom | 63.9 | 95.1 | 80.6 | 47.6 | 71.8 |
| Method | MMLU | SciQA | GSM8K | HumanE. | Avg. |
|---|---|---|---|---|---|
| Regular | 38.9 | 92.4 | 55.1 | 28.2 | 53.7 |
| GradMatch | 38.1 | 92.2 | 53.9 | 26.4 | 52.7 |
| UDS | 40.3 | 93.3 | 56.8 | 28.9 | 54.8 |
| Loom | 41.4 | 93.8 | 54.3 | 27.9 | 54.4 |
| Loom (stratified) | 41.2 | 93.9 | 57.9 | 29.6 | 55.7 |
| Method | MMLU | SciQA | GSM8K | HumanE. | Avg. |
|---|---|---|---|---|---|
| Regular | 46.1 | 88.7 | 63.0 | 36.1 | 58.5 |
| GradMatch | 45.5 | 88.4 | 62.6 | 35.2 | 57.9 |
| UDS | 47.4 | 89.5 | 64.4 | 36.9 | 59.6 |
| Loom | 48.5 | 89.9 | 65.1 | 38.0 | 60.4 |
| Method | GSM8K | MATH500 | SVAMP | Avg. |
|---|---|---|---|---|
| Regular | 56.0 | 17.8 | 71.2 | 48.3 |
| UDS | 58.9 | 18.9 | 72.6 | 50.1 |
| GREATS | 57.0 | 18.1 | 71.8 | 49.0 |
| GradMatch | 55.7 | 17.5 | 70.9 | 48.0 |
| Loom | 58.6 | 19.6 | 73.4 | 50.5 |