Tool use, such as web search, has become a standard capability even in freely available large language models (LLMs). However, existing benchmarks evaluate temporal reasoning mainly in static, non-tool-using settings, which poorly reflect how LLMs perform temporal reasoning in practice. We introduce Time Puzzles, a constraint-based date inference task for evaluating iterative temporal reasoning with tools. Each puzzle combines factual temporal anchors with (cross-cultural) calendar relations and may admit one or multiple valid dates. The puzzles are algorithmically generated, enabling controlled and continual evaluation. Across 13 LLMs, even the best model (GPT-5) achieves only 55.3% accuracy without tools, despite using easily searchable facts. While web search improves performance, models perform substantially better when constraints are rewritten with explicit dates, removing the need for factual lookup. These results reveal a gap in reliable tool use for iterative temporal reasoning.
Figures & tables
Figure 1: Time Puzzles: a simple date inference task requiring iterative tool-aided temporal reasoning.
Figure 2: Average exact match accuracy across solution counts, with and without web search (only on solution counts 1, 3, 5). GPT-4.1/5 run live web search; for open-weight models we re-use the same cited GPT web results.
Model
EM (%)
JI (%)
F1 (%)
# Tks
GPT-5
55.33
58.62
59.52
4126
GPT-5-nano
15.83
19.37
20.31
5224
GPT-4.1
35.17
39.13
40.14
615
GPT-4.1-nano
9.33
13.53
14.84
1098
GPT-oss-20B
10.83
13.51
14.32
5704
DeepSeek-V3.2-R
36.67
39.33
40.08
7862
Table 1: Exact-match accuracy (EM), Jaccard index (JI), F1 score (F1), and output tokens (Tks), averaged over six solution counts, for the 600 generated puzzles with implicit (default) constraints in the tool-less setting.
Figure 3: Average EM accuracy (%) for GPT-5 (left) and GPT-4.1 (right) under different conditions across three solution counts (1, 3, and 5). Numbers annotated on bars are mean values, beside error bars. +Web enables web search. +CI enables Code Interpreter.
Type
Setting
G4.1-Nano
G5-Nano
G4.1
G5
Implicit
Default
9.3
15.8
35.2
55.3
+Reason Inst
8.8
14.2
25.3
47.3
New Runs
11.0
16.8
29.8
50.1
Explicit
Default
41.8
60.2
63.0
84.7
+Reason Inst
35.5
55.5
53.5
77.7
New Runs
42.7
60.2
63.7
80.5
Table 2: EM accuracy (%) averaged over six solution counts under different settings for both explicit and implicit constraint across the four GPT models. Default setting is based on Section 3 . “+Reason Inst” enhances the default setting with detailed step-by-step reasoning instruction. “New Runs” reruns the default setting with completely new puzzles generated from scratch.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Fact Name
Description
Level
ExplicitYearFact
Explicitly states the year (e.g., “The year is 2025”).
Year
DecadeFact
Specifies the decade (e.g., “The 1990s”).
Year
LeapYearFact
States that the year is a leap year.
Year
ChineseZodiacFact
Specifies the Chinese Zodiac animal.
Year
PersonAliveFact
States a famous person (e.g., Steve Jobs) was alive.
Year
USPresidentFact
States a specific US President was in office.
Year
Appendix
Table 3: Taxonomy of temporal facts used in generation, categorized by constraint granularity level.
Model
Variant used (snapshot / HF repo)
Params
Release date
Knowledge cutoff
GPT-5 OpenAI (2025a)
gpt-5-2025-08-07
Not disclosed
2025-08-07
2024-09-30
GPT-5-nano OpenAI (2025a)
gpt-5-nano-2025-08-07
Not disclosed
2025-08-07
2024-05-31
GPT-4.1 OpenAI (2025b)
gpt-4.1-2025-04-14
Not disclosed
2025-04-14
2024-06-01
GPT-4.1-nano OpenAI (2025b)
gpt-4.1-nano-2025-04-14
Not disclosed
2025-04-14
2024-06-01
GPT-oss-20B OpenAI et al. (2025)
openai/gpt-oss-20b
20.9B (3.6B active)
2025-08-05
2024-06
DeepSeek-V3.2-R DeepSeek-AI et al. (2025)
deepseek-reasoner
685B
2025-12-01
Not disclosed
Appendix
Table 4: Model variants and specifications used in our experiments. We identify OpenAI models by their dated snapshot IDs and open-weight models by their Hugging Face repository IDs. Release dates reflect snapshot timestamps or official provider documentation. Note that precise release and knowledge cutoff dates may vary slightly across sources.
Introducing Relay-Bench, an unsaturated, holistic, text-only benchmark that measures LLMs' ability to complete an assortment of tasks from distinct domains in a single prompt. The leading model, GPT-5.5 (xHigh), scores 43.3%. The test set entirely consists of composite problems: groups of single-domain subproblems that are strung together into challenges that require reasoning across multiple domains in combination. Many of these problems then have layers of complexity added through prompt encoding and deliberate context bloat. Domains tested include visual reasoning, coding, math, information extraction (with a focus on web search), problem-solving, general knowledge, and data analysis. No restrictions are imposed outside of the model harness, and models are explicitly encouraged to leverage code-execution, web searches, and all available tools. All problems are composed of two to thirteen subproblems and do not require multi-modal input or output.
Model editing keeps large language models (LLMs) up to date without retraining, but temporal facts expose a limitation of the prevailing locate-and-edit paradigm: an update is not always a replacement. When a fact changes, the new answer should become current while the old answer may remain correct in historical time contexts. Building on this insight, we use causal tracing to show that LLMs already support this distinction via a two-stage internal computation: early MLP layers retrieve a time-agnostic subject representation, and later layers modulate it with temporal context to yield the time-correct answer. Motivated by this finding, we introduce PRISM Edit, which optimizes a single polysemous representation across temporal contexts and leverages the model's inherent modulation pathway to route it to temporally correct predictions without requiring any architectural modification. We evaluate on TimeConflict, a newly introduced temporal editing benchmark, and on temporally augmented CounterFact. PRISM Edit improves multiple core metrics over the best baseline, most notably +23.3 Temporal Consistency (TC) and +33.7 Current Relative-time Score (CRS) on LLaMA-3, while being more than 2x faster. Code and data are publicly available at https://github.com/CheerCHuang/PRISM-Edit.
Large language models (LLMs) often fail to reason under temporal cutoffs: when prompted to answer from the standpoint of an earlier time, they exploit knowledge that became available only later. We study this failure through the lens of ex-ante reasoning, where a model must rely exclusively on information knowable before a cutoff. Through a systematic analysis of prompt-level interventions, we find that temporal leakage is highly sensitive to cutoff formulation and instruction placement: explicit cutoff statements outperform implicit historical framings, and prefix constraints reduce leakage more effectively than suffix constraints. These findings indicate that prompting can steer models into a temporal frame, but does not endow them with the ability to verify whether a response is temporally admissible. We further argue that supervised fine-tuning is insufficient, since ex-ante correctness is not an intrinsic property of an answer, but a relation between the answer and the cutoff. To address this gap, we propose TCFT, a Temporal Critique Fine-Tuning framework that trains models to acquire cutoff-aware temporal verification. Given a query, a cutoff, and a candidate response, TCFT teaches the model to identify post-cutoff leakage, explain temporal boundary violations, and judge temporal admissibility. Experiments with Qwen2.5-7B-Instruct and Qwen2.5-14B-Instruct show that TCFT consistently outperforms prompting and SFT baselines, reducing average leakage by 41.89 and 37.79 percentage points, respectively.
Chenlu Ding, Jiancan Wu, Yanchen Luo +3
University of Science and Technology of China · The Hong Kong Polytechnic University · University of Notre Dame