cs.CLSep 29, 2026

Dating the Model: Hidden Dates in System Prompts Affect LLM Evaluation

Authors: Mario Sanz-Guerrero, Minh Duc Bui, Manuel Mager, Katharina von der Wense

Organizations: Johannes Gutenberg University Mainz, Germany · Universidad Iberoamericana, Mexico · University of Colorado Boulder, USA

Abstract

Reproducibility is essential for scientific research, yet prior work shows that LLM outputs vary with hardware and batching. We identify an overlooked factor: the hidden injection of the current date into system prompts, which users cannot control and which changes every day. Across 9 recent LLMs and 6 datasets spanning multiple-choice QA (MCQA), math reasoning, code generation, and machine translation, performance varies solely with the current date, with deltas of up to 6% on MCQA, 14% on math reasoning, 7% on code generation, and 2.84 BLEU on machine translation. Model rankings also shift, affecting leaderboards. This date effect exceeds other sources of non-determinism, such as batch size and numerical precision. Standard prompting techniques -- chain-of-thought and few-shot prompting -- do not reduce the sensitivity; chain-of-thought even amplifies it. Our findings underscore the need for careful evaluation protocols to ensure reproducibility and fair comparisons in LLM research.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility

    May 19, 2026David Pape, Jonathan Evertz, Lea SchönherrLarge Language Model BenchmarksLLM Inference Optimization