"Don't tune hyperparameters on the test set" is often stated in machine learning textbooks. Violating it is considered a cardinal sin that produces misleadingly optimistic results, corrupts benchmark integrity, and thus can even be interpreted as scientific fraud. Yet evidence suggests that test set hyperparameter tuning does occur in practice, making it all the more important to understand its actual consequences. So how bad is it, really? In this work we question this dogma and put it to an empirical test. We systematically study the magnitude of the performance inflation caused by tuning the hyperparameters on the test set for MNIST-1D, CIFAR-10, and three tasks from the GLUE benchmark. Our experiments show that while the effect is real and significant, it is frequently small relative to other sources of noise. In many cases, we find that tuning on the test set recovers exactly the same model as when tuning on the validation set. Most importantly, we find that the rankings of models remain essentially preserved after tuning on the test set and therefore that consistent test-set tuning may not invalidate benchmarks or model selection. Our results call for a more nuanced view of tuning hyperparameters on the test set, stimulating researchers to openly report test tuning.
Figures & tables
Figure 1: Overfitting sensitivity: the tie-broken win probability of test-tuning, P^(ΔAO>0) , as a function of the tuning budget T . Each line is one architecture, with its 95% percentile-bootstrap CI shaded. A line whose CI lower bound clears the dashed threshold at 0.5 shows statistically significant adaptive overfitting at that budget. MNIST-1D and MRPC clear it at around T=3 , CIFAR-10 at around T=13 . RTE and CoLA are in Appendix A .
Benchmark
ΔAO=0
Percentage
Identical configuration
MNIST-1D
99/240
41.2%
98.0%
CIFAR-10
116/240
48.3%
98.3%
MRPC
31/240
12.9%
90.3%
Table 1: Runs with ΔAO=0 pooled over the eight architectures of each benchmark ( K=30 runs each), as a count and as a percentage of all runs. The last column is the share of runs with ΔAO=0 in which validation-tuning and test-tuning selected the identical configuration. The remainder are distinct configurations that happen to score identically on the test split. RTE and CoLA are in Appendix A .
Figure 2: Overfitting severity: per-run adaptive overfitting gap. Each translucent dot is one of the K=30 paired runs of an architecture; the filled blue circle is the per-architecture mean ΔAO . The red horizontal bar depicts one standard deviation of that architecture’s val-tuned test accuracy, representing run-to-run noise. On nearly every architecture the mean ΔAO sits below the standard deviation, showing that the gap is small relative to natural run-to-run variability. RTE and CoLA are in Appendix A .
Figure 3: Validation-tuned vs. test-tuned mean test accuracy. Markers are per-architecture means with 95% bootstrap confidence intervals; dots are the individual paired runs. The dashed grey line is the identity y=x and the orange line is the ordinary least squares fit through the eight architecture means, with its 95% pairs-bootstrap confidence interval shaded. On every benchmark the means lie close to a line nearly parallel to the identity, so test-tuning acts as a near-uniform upward shift that preserves architecture rankings. RTE and CoLA are in Appendix A .
Val.
Model
Validation-tuned
Test-tuned
ΔAO
Val. Std.
Δ Rank
Rank
(%)
(%)
(%)
MNIST-1D
1
MLP-Giant
64.59
65.56
+0.97
1.76
+0
2
MLP-Huge
64.25
64.96
+0.72
2.46
+0
3
MLP-Larger
63.58
64.75
+1.17
1.97
+0
4
MLP-Large
61.74
62.65
+0.91
1.96
+0
Table 2: Model rankings ordered by validation-tuned accuracy, with the adaptive overfitting difference and the run-to-run standard deviation of validation-tuned accuracy. Several models within a benchmark are close together and have near-similar accuracy. On MNIST-1D and CIFAR-10 no model changes rank when tuned on the test set rather than the validation set. On MRPC exactly one adjacent pair exchanges places, between models separated by far less than the run-to-run standard deviation, so they are statistically indistinguishable under either rule. The top-ranked model is unchanged. RTE and CoLA are in Appendix A .
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
ΔAO=0
Percentage
Identical configuration
RTE
62/240
25.8%
95.2%
CoLA
48/240
20.0%
95.8%
Appendix
Table 3: Runs with ΔAO=0 on RTE and CoLA, pooled over the eight architectures of each dataset ( K=30 runs each). Columns are as in Table 1 .
Figure 4: Overfitting sensitivity on RTE and CoLA: the tie-broken win probability of test-tuning, P^(ΔAO>0) , as a function of the tuning budget T . Each line is one architecture, with its 95% percentile-bootstrap CI shaded. A line whose CI lower bound clears the dashed threshold at 0.5 shows statistically significant adaptive overfitting at that budget. Both datasets clear it at around T=3 . Compare Figure 1 .
Figure 5: Overfitting severity on RTE and CoLA: per-run adaptive overfitting gap. Each translucent dot is one of the K=30 paired runs of an architecture; the filled blue circle is the per-architecture mean ΔAO . The red horizontal bar depicts one standard deviation of that architecture’s val-tuned test accuracy, representing run-to-run noise. Only distilbert on RTE has a mean gap above its standard deviation. Compare Figure 2 .
Figure 6: Validation-tuned vs. test-tuned mean test accuracy on RTE and CoLA. Markers are per-architecture means with 95% bootstrap confidence intervals; dots are the individual paired runs. The dashed grey line is the identity y=x and the orange line is the ordinary least squares fit through the eight architecture means, with its 95% pairs-bootstrap confidence interval shaded. On both datasets the means lie close to a line nearly parallel to the identity. Compare Figure 3 .
Val.
Model
Validation-tuned
Test-tuned
ΔAO
Val. Std.
Δ Rank
Rank
(%)
(%)
(%)
RTE
1
roberta-base
73.62
74.88
+1.26
1.88
+0
2
deberta-v3-small
73.08
74.65
+1.57
1.75
+0
3
xlnet-base
68.53
70.45
+1.92
2.50
+0
4
albert-base
68.35
70.03
+1.68
1.92
+0
Appendix
Table 4: Model rankings on RTE and CoLA, ordered by validation-tuned accuracy. Columns are as in Table 2 . On RTE no model changes rank; on CoLA one adjacent pair exchanges places, between models separated by far less than the run-to-run standard deviation. The top-ranked model is unchanged on both.
Hyperparameter
Distribution
Range
Learning rate
Log-Uniform
[10−4,10−1]
Learning-rate decay rate
Log-Uniform †
[10−5,10−2]
Weight decay
Log-Uniform
[10−6,10−2]
Initialization seed
Uniform integer
[0,231−1]
† Set to zero with probability 0.3 to explicitly sample the no-decay setting.
Appendix
Table 5: Hyperparameter search space for MNIST-1D.
Hyperparameter
Distribution
Range
Learning rate
Log-Uniform
[5×10−3,0.5]
Weight decay
Log-Uniform
[10−5,10−2]
SGD momentum
Uniform
[0.8,0.99]
Training steps
Log-Uniform
[75,000,141,000]
Initialization seed
Uniform integer
[0,231−1]
Appendix
Table 6: Hyperparameter search space for CIFAR-10.
Hyperparameter
Distribution
Range
Learning rate
Log-Uniform
[5×10−6,10−4]
Weight decay
Log-Uniform
[10−5,10−1]
Warmup ratio
Uniform
[0.0,0.2]
Number of epochs
Uniform integer
[2,10]
Initialization seed
Uniform integer
[0,231−1]
Appendix
Table 7: Hyperparameter search space for the GLUE tasks, shared across all three tasks and all eight architectures. The ranges bracket the published fine-tuning grids of the model families in the registry: learning rates of 2 – 5×10−5 (BERT, XLNet), 1 – 3×10−5 (RoBERTa), 1 – 5×10−5 (ALBERT) and 10−4 (ELECTRA-base); weight decays of 0 (HuggingFace run_glue default), 0.01 (BERT, DeBERTa) and 0.1 (RoBERTa); warmup ratios of 6% (RoBERTa) and 10% (BERT); and epoch counts of 3 (BERT) through 10 (RoBERTa).
Architecture
Hidden sizes
Parameters
MLP-Tiny
[25]
1,285
MLP-Mini
[50]
2,560
MLP-Small
[100]
5,110
MLP-Base
[100,100]
15,210
MLP-Large
[100,100,100]
25,310
MLP-Larger
[200,200,200]
90,610
Appendix
Table 8: MLP family evaluated on MNIST-1D. All networks use ReLU activations and residual connections between hidden layers.
#
Model
Paper
Repository
1
ResNet-20
He et al. (2016)
github.com/hysts/pytorch_image_classification
2
ResNet-32
He et al. (2016)
github.com/hysts/pytorch_image_classification
3
ResNet-44
He et al. (2016)
github.com/hysts/pytorch_image_classification
4
ResNet-56
He et al. (2016)
github.com/hysts/pytorch_image_classification
5
Shake-Shake (32d)
Gastaldi (2017)
github.com/hysts/pytorch_image_classification
6
VGG-15 BN
Simonyan and Zisserman (2015)
github.com/hysts/pytorch_image_classification
Appendix
Table 9: Architectures evaluated on CIFAR-10. Each model is ported from its reference implementation without architectural modifications.
#
Model
Paper
Parameters
Hugging Face checkpoint
1
ALBERT-base-v2
Lan et al. (2020)
11.8 M
albert/albert-base-v2
2
ELECTRA-small
Clark et al. (2020)
14.0 M
google/electra-small-discriminator
3
MobileBERT
Sun et al. (2020)
25.3 M
google/mobilebert-uncased
4
DistilBERT-base
Sanh et al. (2019)
66.4 M
distilbert/distilbert-base-uncased
5
BERT-base
Devlin et al. (2019)
109.5 M
google-bert/bert-base-uncased
6
XLNet-base
Yang et al. (2019)
110.0 M
xlnet/xlnet-base-cased
Appendix
Table 10: Pretrained transformers fine-tuned on each GLUE task. All checkpoints are public Hugging Face models, used without architectural modification; each trial loads the checkpoint afresh and attaches a randomly initialised classification head. The same eight models are evaluated on all three tasks.
Figure 7: Upper-bound reference per architecture across the five benchmarks, ordered by mean validation-tuned test accuracy. Blue circles are the validation-tuned mean test accuracy and orange squares the test-tuned mean test accuracy (both with 95% bootstrap CIs over the K=30 runs). Orange triangles are the upper bound: the best test accuracy among T=30 models trained directly on the same bootstrap test fold. The vertical distance between the blue marker and the orange triangle is the headroom available to any procedure that exploits the test set; the much smaller distance between the blue circle and orange square is the share that random search over T=30 trials actually captures.
Benchmark
Upper bound (%)
Headroom (%)
Captured (%)
MNIST-1D
100.00
35.4 – 46.0
2.0 – 3.7
CIFAR-10
100.00
4.6 – 7.7
1.2 – 2.4
RTE
99.80 – 100.00
26.4 – 40.9
4.4 – 6.9
MRPC
99.06 – 100.00
13.6 – 18.7
6.9 – 9.6
CoLA
98.87 – 100.00
13.8 – 19.7
1.0 – 5.5
Appendix
Table 11: Ranges over the eight architectures of each benchmark. Upper bound is the best test accuracy among T=30 models trained and tuned directly on the bootstrap test fold. Headroom is the distance from the validation-tuned mean to that bound, and captured is the share of it that test-tuning actually realises, (A(λtest∗)−A(λval∗))/headroom .
Figure 8: MLP-Giant rerun with T=100 trials per run on larger 10,000/10,000/10,000 MNIST-1D splits. Left: per-run ΔAO across the K=30 paired runs; the blue dot is the mean and the red horizontal bar marks the standard deviation of val-tuned accuracy. Right: P^(ΔAO>0) as a function of hyperparameter tuning budget T , with the 95% percentile-bootstrap CI; the dashed grey reference at 0.5 is the significance threshold. The mean gap stays small relative to the val-tuned noise and P^(ΔAO>0) remains significant ( CIlower>0.5 ) across the full budget range.
Model
Val-tuned (%)
Test-tuned (%)
CIFAR-10.1-tuned (%)
on CIFAR-10.1
on CIFAR-10.1
on CIFAR-10.1
shake_shake_32d
88.79
88.82
89.11
resnet9
86.95
86.93
87.41
mobilenetv2
87.11
87.06
87.29
resnet_basic_56
86.07
86.11
86.39
resnet_basic_44
85.95
85.78
86.26
Appendix
Table 12: Per-architecture mean CIFAR-10.1 accuracy under three selection rules. Val-tuned and Test-tuned pick the trial with the highest accuracy on the bootstrap validation or test fold, respectively, and the picked configuration is then re-evaluated on CIFAR-10.1; CIFAR-10.1-tuned picks the trial with the highest CIFAR-10.1 accuracy directly. “Pooled” is the mean across the eight architectures. The val-tuned and test-tuned columns are indistinguishable on CIFAR-10.1 (pooled gap −0.04% , within run-to-run noise), consistent with the small ΔAO≈0.12% already observed on the bootstrap test fold. CIFAR-10.1-tuned wins by +0.43% pooled — a small but consistent margin that the bootstrap-test-tuned configuration does not capture.
Figure 9: MNIST-1D: empirical distribution of winning hyperparameter values for the validation-tuned and test-tuned selection rules. 30.8% of trials had lr_decay=0 and are excluded from the lr_decay panel. All three HPs have their modes well inside the bounds with no edge clustering.
Figure 10: CIFAR-10: empirical distribution of winning hyperparameter values for the val-tuned and test-tuned selection rules. lr , weight_decay , and momentum have their modes well inside the bounds; only max_steps clusters at the upper bound of 141,000 , suggesting the optimum lies above this cap. Extending it would raise absolute accuracies but is unlikely to change ΔAO , since both selection rules benefit equally from longer training.
Figure 11: RTE: empirical distribution of winning hyperparameter values for the val-tuned and test-tuned selection rules, over the 240 winners of each rule ( 8 architectures ×K=30 runs). num_epochs is binned on the integers 2 – 10 .
Figure 12: MRPC: empirical distribution of winning hyperparameter values for the val-tuned and test-tuned selection rules, over the 240 winners of each rule.
Figure 13: CoLA: empirical distribution of winning hyperparameter values for the val-tuned and test-tuned selection rules, over the 240 winners of each rule. As on the other two GLUE tasks, num_epochs rises towards the upper bound of the range.
Post-training hyperparameter selection is a critical step in the deployment of modern artificial intelligence systems, given the need to tune degrees of freedom of pre-trained models such as inference-time parameters, implementation-level settings, and thresholds driving decision rules. Despite its practical importance, hyperparameter selection is typically performed using best-effort empirical methods such as grid search or Bayesian optimization, which provide no formal statistical guarantees on reliability or safety. This monograph, intended for an audience of signal processing and machine learning researchers, presents a unified statistical framework for reliable post-training hyperparameter selection, centered on the learn-then-test (LTT) paradigm. LTT formulates the hyperparameter selection problem as multiple hypothesis testing over a candidate set of hyperparameters. The framework enables the choice of hyperparameters that provably satisfy application-specific reliability requirements---such as bounds on average risk, quantile risk, or information-theoretic constraints---with explicit, finite-sample control of error probabilities. The supporting statistical machinery, namely p-values, e-values, and concentration inequalities, is developed from first principles.
Amirmohammad Farzaneh, Osvaldo Simeone
Institute for Intelligent Networked Systems (INSI) Northeastern University London
Every supervised fine-tuning run forces the same chain of decisions, such as learning rate, batch size, LoRA or full fine-tuning, how many epochs, which optimiser, and what data to feed the model. Each of these is typically rediscovered from scratch for every new model and dataset. Here we measure them under one instrument: a sweep that varies one lever at a time, and spans dense and mixture-of-experts models in two families (Qwen3 and Llama), on four real-world customer SFT datasets, for both LoRA and full fine-tuning. These datasets give a controlled testbed: each task carries an evaluation built with the customer, and its training data is produced by iterative supervised fine-tuning that refines model outputs until they pass that evaluation, so the supervised target is internally consistent and the task judge we report against is the criterion the data was built to satisfy. We ask how the optimal learning rate and batch size move with model scale, family, and data, and whether one selection rule transfers across them; what LoRA trades against full fine-tuning, and how its rank and alpha set what the adapter can learn; whether validation loss (or other metrics, such as loss landscape flatness) faithfully ranks downstream quality; whether post-training gains scale with model size and data volume, on a model ladder extended through mixtures-of-experts to 235B parameters; how many epochs to train before general instruction-following erodes; and whether a geometry-aware optimiser improves on AdamW. Each recommendation is paired with a measure of its uncertainty.
Charles O'Neill, Mudith Jayasekara, Harry Partridge
Artificial intelligence benchmarks are an important mechanism for measuring model progress and guiding deployment decisions. However, benchmarks quickly "saturate", making it difficult to differentiate models and diminishing their long-term value. In this study, we define benchmark saturation and analyze it across 60 language model benchmarks using 14 properties that relate to saturation. We find that nearly half of the our benchmarks exhibit saturation, with rates increasing with age. Further, we find that resilience to saturation is impacted by expert-curation, not by public test data. Our results suggest that design choices can extend benchmark longevity and inform more durable evaluation approaches.
Mubashara Akhtar, Anka Reuel, Prajna Soni +34
ETH Zurich · ETH AI Center · Stanford University +23