cs.AIOct 5, 2026

Do Time-Series QA Systems Read the Time Series? Evidence Use and Reasoning Reliability

Authors: Zhuomin Chen, Jingchao Ni, Xu Zheng, Janki Bhimani, Mo Sha, Wei Cheng, Dongsheng Luo

Organizations: Florida International University · University of Houston · NEC Laboratories America · Singapore Management University

Abstract

In recent years, time-series question answering (QA) systems have made significant progress. However, generating a correct answer does not show whether retaining the supplied numerical series improves task performance, nor whether the prediction is sensitive to changes in that input. While some systems provide rationales, answer accuracy also does not show whether their numerical claims are grounded in the supplied series or whether the stated inference is valid. In this work, we focus on evaluating four time-series QA systems: TimeOmni-1, ChatTS, TimeOmni-VL, and Time-MQA. First, for three systems with released evaluation data, we reproduce their reported results and compare the performance of the systems with their backbones. Then, we introduce a benchmark named COMMON-TSQA, which collects public evaluation datasets from existing time-series benchmarks and unifies their sample representation, task definitions, and answer schemas, while evaluating each system through its own interface under common evaluation criteria. The evaluation uses the original condition and six interventions while keeping the question and target fixed. Our analysis shows that aggregate performance alone can obscure how systems use numerical evidence. Similar task-level scores can arise despite substantial changes in individual predictions. Some interventions induce simple fallback behavior rather than preserved task ability. We also evaluate rationales for factual grounding, inference validity, and consistency with the final answer. We find that rationales often contain time-series claims unsupported by the input. Moreover, the rationale audit shows that agreement between a rationale and its final answer can coexist with incorrect numerical descriptions or invalid intermediate inferences.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ARFBench: Benchmarking Time Series Question Answering Ability for Software Incident Response

    Apr 23, 2026Stephan Xie, Ben Cohen, Mononito Goswami +6Time Series ReasoningTime Series QA

  2. Towards Verifiable Agentic Data Science: Solving Irregular TSQA Via Tool-Grounded Reasoning

    Jun 13, 2026Sanhorn Chen, Xiaoyang Chen, Boyu Liu +1Irregular Time-Series ModelingLLM Evaluation

  3. TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering

    May 23, 2026Liying Han, Kang Yang, Oliver Wang +9LLM EvaluationTime Series Reasoning