As frontier AI systems are pretrained on web-scale data, test set contamination has become a critical concern for accurately assessing their capabilities. While research has thoroughly investigated the impact of test set contamination on discriminative evaluations like multiple-choice question-answering, comparatively little research has studied the impact of test set contamination on generative evaluations. In this work, we quantitatively assess the effect of test set contamination on generative evaluations through the language model lifecycle. We pretrain language models on mixtures of web data and the MATH benchmark, sweeping model sizes and number of test set replicas contaminating the pretraining corpus; performance improves with contamination and model size. Using scaling laws, we make a surprising discovery: including even a single test set replica enables models to achieve lower loss than the irreducible error of training on the uncontaminated corpus. We then study further training: overtraining with fresh data reduces the effects of contamination, whereas supervised finetuning on the training set can either increase or decrease performance on test data, depending on the amount of pretraining contamination. Finally, at inference, we identify factors that modulate memorization: high sampling temperatures mitigate contamination effects, and longer solutions are exponentially more difficult to memorize than shorter ones, presenting a contrast with discriminative evaluations, where solutions are only a few tokens in length. By characterizing how generation and memorization interact, we highlight a new layer of complexity for trustworthy evaluation of AI systems.
Figures & tables
Figure 1: Performance on Generative Benchmarks Increases with Test Set Contamination and Model Size. We pretrained compute-optimal language models ( 34 M– 344 M parameters) on corpora containing 0 – 3162 replicas of the MATH test set, then evaluated with greedy decoding. As contamination grows, Math Verify scores rise (top) and test-set cross entropies fall (bottom), with sharp improvement around 100 replicas. The ratio of test-set loss at R replicas to loss at 0 replicas grows with model size, so larger models benefit more from contamination at any replica count.
Figure 2: Scaling Laws Suggest Including A Single Test Set Replica Achieves Lower Loss Than the Irreducible Error of the Uncontaminated Corpus. Top: For each scaling series contaminated with R replicas of the MATH test set, we fit scaling laws L(C,R)=E(R)+C0(R)⋅C−α(R) , where C≈6ND is pretraining compute. Almost all contaminated models achieve lower cross entropy on the test set than the irreducible error of training on uncontaminated data (horizontal purple line). Bottom: Increasing test set contamination reduces the irreducible error E(R) from 3.594 at R=0 to 0.0347 at R=316 . Larger values of R also increase the compute prefactor and compute exponent. The functional form achieves average fitting error <10−2 for all R .
Table 1: Contamination-Driven Performance Is Memorization, Not Generalization. When MATH test problems are rephrased (same numbers, different wording) or perturbed (same wording, different numbers), model performance collapses to baseline regardless of contamination level or model size, confirming that gains from contamination are not attributable to general mathematical reasoning. Results were consistent across all model sizes.
Figure 3: Overtraining with Fresh Data Mitigates Contamination. Consistent with discriminative evaluations ( Bordt et al., 2025 ) , overtraining (training longer than Chinchilla compute-optimal) on new fresh data interacts with contamination: increasing overtraining decreases cross entropy on the MATH test set for uncontaminated models, but increases cross entropy for contaminated models. This suggests that while fresh data generally improves performance, it dilutes the “dose,” or proportion of contaminated pretraining tokens, thereby lessening the effect of the contamination. The crossover point shifts with model size, falling from 32 test set replicas for 34M to 1 replica for 93M models, indicating larger models lose their contamination advantage more readily when overtrained.
Figure 4: Supervised Finetuning on the Train Set Has Opposing Effects, Depending on Pretraining Contamination. For little-to-no contamination ( <10 test set replicas), supervised finetuning (SFT) on the MATH train set decreases loss on the test set. For models pretrained with more contamination ( >10 replicas), SFT on the train set increases loss on the test set.
Figure 5: Sampling Temperature Degrades Performance, Particularly for Contaminated Models. As sampling temperature increases, Math Verify scores drop, falling from near 100% to under 1% in many configurations. The drop is disproportionately larger for highly contaminated models: increasing temperature from 0 to 1 reduces performance by a factor of ∼2 at low contamination ( ≤10 replicas), whereas it reduces performance by a factor of up to 40 at high contamination.
Figure 6: Performance Declines with Increasing Solution Token Length. Math Verify Scores decrease with increasing solution length, modulated by temperature. At high rates of contamination and higher temperatures, Math Verify scores fall exponentially with the solution length. At lower temperatures, and high rates of contamination, solution length has almost no effect. Models exposed to less contamination also exhibit declining scores with longer solution lengths, but the dropoff is concave up and occurs much more slowly.
Figure 7: Solution Length and Temperature Together Govern How Quickly Memorization Decoheres. Top: Most models and contamination pairs exhibit scaling laws as a function of sequence length (Eqn. 2 ). Bottom: The cumulative probability of generating a memorized solution lives in one of three regimes: (1) Exponentially Fast Decoherence, (2) Brittle Memorization, and (3) Deterministic Lock-In. Sampling temperature shifts the scaling exponent, and can shift a model between regimes.
Figure 8: Regimes of Memorization. Phase diagrams illustrating the maximum sustainable solution length T for a fixed survival probability P(T)=0.01 . The three regimes are (1) Exponentially Fast Decoherence, (2) Brittle Memorization, and (3) Deterministic Lock-In.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Parameters
Num. Layers
Hidden Size
34M
3
96
62M
5
160
93M
6
224
153M
9
320
344M
14
576
Appendix
Table 2: Model architecture configurations following Qwen 3 scaling patterns.
Figure 9: Deviations from In-Context / Sequence Scaling Laws. Two groups of models exhibit increasing negative log likelihoods with increasing token index: smaller uncontaminated models and larger massively contaminated models.
Figure 10: Fit Parameters for In-Context / Sequential Scaling Laws by Model Size and Number of Test Set Replicas.
Figure 11: Math Verify Score Is Correlated with Pretraining Loss. Math Verify scores correlate strongly with the cross entropy loss achieved on the MATH Test Set during training, where differences in these graphs are attributable to increased repetition on the benchmark test set. The correlation is significantly weaker for high temperatures and falls to nearly 0 for temperatures above 1.0 .
Benchmark datasets are critical for reproducible, reliable, and discriminative evaluation of LLMs. However, recent studies reveal that many benchmark datasets are included in pretraining corpora, i.e., contaminated, which diminishes their value as reliable measures of model generalization. In this paper, we argue that benchmark datasets should be contamination-resistant, i.e., unlearnable, but support inference. To accomplish this, we first highlight the wide prevalence of benchmark dataset contamination and outline the properties of contamination-resistant datasets. Second, we highlight how the asymmetry between the inference and training pipelines in the Transformer architecture can be leveraged to support contamination-resistance. Third, we outline mathematical advancements to make these datasets interoperable across various LLM architectures. Based on the above, we call on the community to ensure the reliability of LLM benchmarking by: (i) advancing novel contamination-resistant methodologies, (ii) developing supporting methods and platforms, and (iii) adopting contamination-resistant benchmarks into existing evaluation pipelines.
Ali Al-Lawati, Jason Lucas, Dongwon Lee +1
The Pennsylvania State University, University Park, PA, USA.
As artificial intelligence (AI)-generated content proliferates, models are increasingly trained on their own outputs, risking progressive degradation or collapse. In this article, we provide the first positive, rigorous theoretical results, to the best of our knowledge, showing that under model-agnostic mild conditions, the model converges to the true data-generating distribution. The convergence rate is the minimum of the model's intrinsic rate and the fraction of real data at each training iteration, revealing a phase transition between data-limited and model-limited regimes. We further show that, for biased real data, correcting the bias prevents the persistence and amplification of early bias over training iteration. Extensive experiments across simulations, real images and texts validate our theoretical framework, establishing quantitative conditions for long-term AI stability in contaminated environments.
Kevin Wang, Hongqian Niu, Didong Li
Department of Biostatistics, University of North Carolina at Chapel Hill
Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorized. \textbf{Contamination mitigation evaluation} intervenes in the decoding process to suppress memorization and restore a contaminated model's genuine capability, but its prevailing metric, the \textbf{G-AP} (\textbf{G}ap of \textbf{A}ggregate \textbf{P}erformance), is flawed. Discrete correct/incorrect readouts cannot characterize per-question performance, averaging before differencing lets over- and under-suppression cancel out, and uniform per-question weighting invites strategies to push solve probabilities onto the clean model's high-frequency values. We propose \textbf{SA-PPG} (\textbf{S}tratified \textbf{A}ggregate of \textbf{P}er-question \textbf{P}robability \textbf{G}aps): estimate each question's solve probability by sampling, difference it against the clean model per question, and aggregate within groups defined by the clean model's solve probability. Existing mitigation strategies first estimate where contamination lies and then operate on the estimate, so they are only as correct as the estimate. \textbf{RailCap} instead judges contamination during generation: whenever a sample falls back onto the greedy trajectory, the next trajectory token is capped to the runner-up, accumulating suppression until the response distribution becomes sufficiently dispersed. Across multiple contaminated models and benchmarks, SA-PPG reveals that prior strategies' restoration is substantially overestimated, while RailCap attains the lowest SA-PPG.