Reproducibility is essential for scientific research, yet prior work shows that LLM outputs vary with hardware and batching. We identify an overlooked factor: the hidden injection of the current date into system prompts, which users cannot control and which changes every day. Across 9 recent LLMs and 6 datasets spanning multiple-choice QA (MCQA), math reasoning, code generation, and machine translation, performance varies solely with the current date, with deltas of up to 6% on MCQA, 14% on math reasoning, 7% on code generation, and 2.84 BLEU on machine translation. Model rankings also shift, affecting leaderboards. This date effect exceeds other sources of non-determinism, such as batch size and numerical precision. Standard prompting techniques -- chain-of-thought and few-shot prompting -- do not reduce the sensitivity; chain-of-thought even amplifies it. Our findings underscore the need for careful evaluation protocols to ensure reproducibility and fair comparisons in LLM research.
Figures & tables
Figure 1: The performance of LLMs on standard benchmarks varies solely by changing the current date, which is usually hidden from users in the system prompt.
Model
MMLU
GPQA
ARC-C
Avg.
Llama 3.1 (8B)
2.11
4.53
2.01
2.88
Llama 3.1 (70B)
2.46
3.54
0.67
2.22
Gemma 3 (4B)
1.75
3.03
1.87
2.22
Gemma 3 (27B)
1.38
4.04
1.01
2.14
Qwen3 (4B)
2.46
2.53
1.34
2.11
Qwen3-Next (80B)
2.11
2.43
1.01
1.85
Table 1: Difference in accuracy (delta) from the worst to the best date in 2024 across models and datasets.
Figure 2: Accuracy (top) and ECE (bottom) across different dates in 2024 for the top-5 models on MMLU.
GSM8K
HumanEval
WMT (en → de)
WMT (en → fi)
WMT (en → cs)
Model
Δ Acc.
Δ pass@1
Δ BLEU
Δ chrF
Δ BLEU
Δ chrF
Δ BLEU
Δ chrF
Llama 3.1 (8B)
9.85
7.32
1.66
1.13
1.79
1.52
1.82
1.42
Llama 3.1 (70B)
7.58
5.49
1.58
1.07
1.71
1.46
1.76
1.36
Gemma 3 (4B)
8.33
6.71
1.57
1.15
1.64
1.11
1.68
1.22
Gemma 3 (27B)
3.03
3.66
1.58
0.96
1.38
0.91
1.90
1.06
Qwen3 (4B)
6.06
2.44
1.34
0.95
1.39
1.38
1.88
1.30
Table 2: Worst-to-best date delta in 2024 across models on open-ended generation tasks: accuracy on GSM8K, pass@1 on HumanEval, and BLEU and chrF on 3 WMT language pairs (English to German, Finnish, and Czech).
Pattern
N
Median
Range
Same dir.
Trend over the year
27
0.02
[ − 0.40, 0.53]
–
Across datasets
27
− 0.03
[ − 0.16, 0.15]
51%
Across models
108
0.01
[ − 0.17, 0.25]
51%
Table 3: Correlation between the date and accuracy (trend over the year; Spearman), and between the accuracies across dates of the same model on two datasets or of two models on the same dataset (Pearson), on MCQA. N : number of model–dataset combinations (trend) or compared pairs. Same dir. : share of dates where both accuracies are above or both below their yearly average (50% = chance).
Figure 3: Accuracy deviation across different dates for GPT-5.1. The 0.0 line represents the average accuracy over the week for each dataset.
Figure 4: Comparison of accuracy across different dates in 2024 for Llama 3.1 (8B) on MMLU using zero-shot prompting vs. chain-of-thought prompting.
Model
MMLU
GPQA
ARC-C
Avg.
Llama 3.1 (8B)
2.81
5.05
2.35
3.40
Llama 3.1 (70B)
1.40
5.56
1.01
2.66
Gemma 3 (4B)
2.46
3.54
1.34
2.45
Gemma 3 (27B)
0.70
1.52
0.67
0.96
Qwen3 (4B)
1.75
3.03
1.43
2.07
Qwen3-Next (80B)
1.75
3.54
1.34
2.21
Table 4: Difference in accuracy (delta) from the worst to the best date in 2024 across models and datasets using 5-shot prompting.
Source of Variation
Acc.
ECE
Current Date [Jan 1 – Dec 31]
0.78 %
2.56 %
Batch Size [1, 2, 4, 8, …, 128]
0.61 %
2.00 %
Num. Precision [BF16, FP16, FP32]
0.52 %
2.54 %
GPU Model [A100, A40, RTX4090]
0.48 %
1.60 %
Option Order [5 random permutations]
0.75 %
3.12 %
System Instruction [6 different wordings]
0.78 %
3.54 %
Table 5: Coefficient of variation (CV) of accuracy and ECE from different sources of non-determinism using Llama 3.1 (8B) on MMLU.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Response of Qwen3 called through OpenRouter, with an empty system prompt, no tool use, and no internet access. The model reports the current date, showing that the provider injects it server-side.
Figure 6: Prompt used for multiple-choice questions. We extract the probabilities of the tokens after the arrow ( → ). “ X ” denotes the option label ( A / B / C / D ).
Figure 7: Prompt used for chain-of-thought MCQA. The model generates a reasoning chain ending in “ The answer is X. ” (where X = A / B / C / D ); we parse the final option label X from this string.
Figure 8: Prompt used for GSM8K. The model generates a step-by-step solution ending with “ Answer: {number} ”, from which we parse the final number.
Figure 9: Prompt used for HumanEval. The model completes the given Python function, and we run it against the unit tests to compute pass@1.
Figure 10: Prompt used for machine translation. {target_language} is German, Finnish, or Czech, depending on the language pair.
Dataset
# Questions
% Time-Dependent
MMLU
14042
0.0 %
GPQA
198
0.0 %
ARC-Challenge
1172
0.0 %
GSM8K
1319
0.0 %
HumanEval
164
0.0 %
WMT en → de
2999
0.0 %
Appendix
Table 6: Number and percentage of time-dependent questions in each dataset, as determined by GPT-OSS (120B).
Model
Consistent
Inconsistent
Llama 3.1 (8B)
83.71 %
41.85 %
Llama 3.1 (70B)
92.13 %
47.59 %
Gemma 3 (4B)
99.49 %
88.86 %
Gemma 3 (27B)
99.92 %
80.48 %
Qwen3 (4B)
98.96 %
66.69 %
Qwen3-Next (80B)
98.82 %
64.38 %
Appendix
Table 7: Average confidence (token probability of the selected answer) for consistent and inconsistent predictions across all dates in 2024 on MMLU.
Model
MMLU
GPQA
ARC-C
Avg.
Llama 3.1 (8B)
2.38
4.17
1.95
2.83
Llama 3.1 (70B)
3.35
4.61
1.24
3.07
Gemma 3 (4B)
2.05
2.98
1.78
2.27
Gemma 3 (27B)
1.49
3.23
0.79
1.84
Qwen3 (4B)
2.41
2.93
1.52
2.29
Qwen3-Next (80B)
2.47
3.39
0.75
2.20
Appendix
Table 8: Difference in ECE (delta) from the worst to the best date in 2024 across models and datasets.
MMLU
GPQA
ARC-C
Model
Acc.
ECE
Acc.
ECE
Acc.
ECE
Llama 3.1 (8B)
67.0
14.5
29.3
28.4
81.9
9.4
Llama 3.1 (70B)
79.4
11.2
40.1
33.6
92.6
4.8
Gemma 3 (4B)
54.4
44.4
29.3
67.1
78.6
21.1
Gemma 3 (27B)
75.0
24.2
33.8
62.1
93.9
5.7
Qwen3 (4B)
73.7
24.3
45.0
46.1
89.0
10.5
Appendix
Table 9: Accuracy and ECE for all models and datasets when the current date is not injected in the system prompt.
Figure 11: Accuracy and ECE across different dates in 2024 for all models on MMLU.
Figure 12: Accuracy and ECE across different dates in 2024 for all models on GPQA.
Figure 13: Accuracy and ECE across different dates in 2024 for all models on ARC-Challenge.
Figure 14: Impact of different user locations in the system prompt on Llama 3.1 (8B) performance on MMLU.