cs.LGOct 5, 2026

On Hyperparameter Tuning on the Test Set

Authors: Matteo Fregonara, Tom Viering, Jan van Gemert

Organizations: Delft University of Technology

Abstract

"Don't tune hyperparameters on the test set" is often stated in machine learning textbooks. Violating it is considered a cardinal sin that produces misleadingly optimistic results, corrupts benchmark integrity, and thus can even be interpreted as scientific fraud. Yet evidence suggests that test set hyperparameter tuning does occur in practice, making it all the more important to understand its actual consequences. So how bad is it, really? In this work we question this dogma and put it to an empirical test. We systematically study the magnitude of the performance inflation caused by tuning the hyperparameters on the test set for MNIST-1D, CIFAR-10, and three tasks from the GLUE benchmark. Our experiments show that while the effect is real and significant, it is frequently small relative to other sources of noise. In many cases, we find that tuning on the test set recovers exactly the same model as when tuning on the validation set. Most importantly, we find that the rankings of models remain essentially preserved after tuning on the test set and therefore that consistent test-set tuning may not invalidate benchmarks or model selection. Our results call for a more nuanced view of tuning hyperparameters on the test set, stimulating researchers to openly report test tuning.

Figures & tables

Appendix figures & tables20 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Statistically Valid Post-Training Hyperparameter Selection: From Tuning to Guarantees

    Jun 24, 2026Amirmohammad Farzaneh, Osvaldo SimeoneTwo-Sample TestingModel Selection

  2. Post-Training Science for Supervised Fine-Tuning

    Sep 1, 2026Charles O'Neill, Mudith Jayasekara, Harry PartridgeModel Fine-TuningBatch

  3. When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

    Feb 18, 2026Mubashara Akhtar, Anka Reuel, Prajna Soni +34Artificial Intelligence BenchmarksLarge Language Model Benchmarks