cs.CLJan 12, 2026

Measuring Iterative Temporal Reasoning with Time Puzzles

Authors: Zhengxiang Wang, Zeyu Dong

Organizations: Department of Applied Math and Statistics Stony Brook University

Abstract

Tool use, such as web search, has become a standard capability even in freely available large language models (LLMs). However, existing benchmarks evaluate temporal reasoning mainly in static, non-tool-using settings, which poorly reflect how LLMs perform temporal reasoning in practice. We introduce Time Puzzles, a constraint-based date inference task for evaluating iterative temporal reasoning with tools. Each puzzle combines factual temporal anchors with (cross-cultural) calendar relations and may admit one or multiple valid dates. The puzzles are algorithmically generated, enabling controlled and continual evaluation. Across 13 LLMs, even the best model (GPT-5) achieves only 55.3% accuracy without tools, despite using easily searchable facts. While web search improves performance, models perform substantially better when constraints are rewritten with explicit dates, removing the need for factual lookup. These results reveal a gap in reliable tool use for iterative temporal reasoning.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. PRISM Edit: One Vector for All Temporal Answers

    Jul 13, 2026Chen Huang, Qi Zheng, Ruiqin Zheng +2Multimodal Knowledge EditingAnswer

  2. Teaching Large Language Models When Not to Know: Learning Temporal Critique for Ex-Ante Reasoning

    May 14, 2026Chenlu Ding, Jiancan Wu, Yanchen Luo +3LLM Reasoning StrategiesContinual Fine-Tuning