Do Time-Series QA Systems Read the Time Series? Evidence Use and Reasoning Reliability
Organizations: Florida International University · University of Houston · NEC Laboratories America · Singapore Management University
Abstract
In recent years, time-series question answering (QA) systems have made significant progress. However, generating a correct answer does not show whether retaining the supplied numerical series improves task performance, nor whether the prediction is sensitive to changes in that input. While some systems provide rationales, answer accuracy also does not show whether their numerical claims are grounded in the supplied series or whether the stated inference is valid. In this work, we focus on evaluating four time-series QA systems: TimeOmni-1, ChatTS, TimeOmni-VL, and Time-MQA. First, for three systems with released evaluation data, we reproduce their reported results and compare the performance of the systems with their backbones. Then, we introduce a benchmark named COMMON-TSQA, which collects public evaluation datasets from existing time-series benchmarks and unifies their sample representation, task definitions, and answer schemas, while evaluating each system through its own interface under common evaluation criteria. The evaluation uses the original condition and six interventions while keeping the question and target fixed. Our analysis shows that aggregate performance alone can obscure how systems use numerical evidence. Similar task-level scores can arise despite substantial changes in individual predictions. Some interventions induce simple fallback behavior rather than preserved task ability. We also evaluate rationales for factual grounding, inference validity, and consistency with the final answer. We find that rationales often contain time-series claims unsupported by the input. Moreover, the rationale audit shows that agreement between a rationale and its final answer can coexist with incorrect numerical descriptions or invalid intermediate inferences.
Figures & tables
| Task | Metric | Model | Original | Shuffle | Flat | Noise | Swap | Answer-only | Zero |
| Forecasting | MASE | TimeOmni-1 | 1.739 | 1.929 | 1.321 | 2.060 | 26.848 | 17.418 | 7.235 |
| ChatTS | – | – | – | – | – | – | – | ||
| Time-MQA | 2.258 | 2.287 | 2.106 | 2.260 | 23.380 | 38.241 | 12.706 | ||
| TimeOmni-VL | 2.650 | 3.001 | 2.604 | 3.291 | 36.109 | 38.546 | 13.221 | ||
| Anomaly detection | Accuracy (%) | TimeOmni-1 | 60.70 | 55.44 | 45.61 | 46.32 | 45.96 | 45.61 | 45.61 |
| ChatTS | 62.46 | 68.88 | 45.61 | 58.60 | 54.39 | 54.39 | 45.61 |
| System | C W | W C | Changed |
| (%) | (%) | (%) | |
| TimeOmni-1 | 24.11 | 16.34 | 48.22 |
| ChatTS | 14.13 | 12.84 | 30.50 |
| Time-MQA | 17.62 | 11.88 | 34.65 |
| TimeOmni-VL | 19.69 | 20.38 | 47.95 |
| Factual grounding | Inference validity | Conclusion consistency | |
| TimeOmni-1 | 2.1 | 27.3 | 64.0 |
| TimeOmni-VL | 11.2 | 34.3 | 92.2 |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| System | Backbone | TS input | Training | Evaluation data |
| TimeOmni-1 | Qwen2.5-7B | Text | CoT-SFT GRPO | TSR-Suite |
| ChatTS | Qwen2.5-14B | TS encoder | Alignment SFT | Eval A/B |
| TimeOmni-VL | BAGEL-7B | TS image | CoT post-training | TSR-Suite (OOD) |
| Time-MQA | Qwen2.5-7B | Text + context | Continual PT (LoRA) | TSQA |
| System | Split | Task/metric | Reproduced | Paper | Difference |
| TimeOmni-1 | ID | Scenario ACC | 88.50 | 90.70 | |
| ID | Causality ACC | 67.60 | 69.30 | ||
| ID | Forecasting MAE | 15.95 | 14.30 | ||
| ID | Decision ACC | 47.90 | 47.90 | ||
| OOD | Scenario ACC | 84.80 | 87.70 | ||
| OOD | Causality ACC | 62.30 | 64.00 |
| System / backbone | Split | Task / metric | System | Backbone | Improvement |
| TimeOmni-1 / Qwen2.5-7B | ID | Scenario ACC | 88.50 | 41.00 | +47.50 |
| Causality ACC | 67.63 | 18.75 | +48.88 | ||
| Forecasting MAE | 15.95 | 23.27 | +7.32 | ||
| Decision ACC | 47.87 | 23.94 | +23.94 | ||
| OOD | Scenario ACC | 84.76 | 43.05 | +41.71 | |
| Causality ACC | 62.25 | 21.13 | +41.12 |
| Task | Data sources | Samples | Domains |
| Forecasting | GIFT-Eval; SciTS; TSR-Suite (OOD); TIME | 605 | Energy; environment; finance; healthcare; information technology; transportation; climate |
| Anomaly detection | TSB-AD-U | 285 | Environment; healthcare; industry; information technology; synthetic |
| Pattern recognition | TSAQA; TSRBench | 623 | Energy; environment; finance; healthcare; information technology; synthetic; transportation |
| Cross-series comparison | TSAQA; CaTS-Bench | 517 | Energy; environment; finance; healthcare; information technology; transportation; climate |
| Causal diagnosis | FactoryBench; SenTSR-Bench; TSR-Suite (OOD) | 395 | Industry; hydrology |
| Decision intervention | SenTSR-Bench; TSR-Suite (OOD) | 175 | Industry; energy |
| Source | Access link |
| GIFT-Eval | https://huggingface.co/datasets/Salesforce/GiftEval |
| TIME | https://huggingface.co/datasets/Real-TSF/TIME |
| SciTS | https://huggingface.co/datasets/OpenTSLab/SciTS |
| TSB-AD-U | https://www.thedatum.org/datasets/TSB-AD-U.zip |
| TSAQA | https://huggingface.co/datasets/TSAQA/TSAQA-Benchmark |
| TSRBench | https://huggingface.co/datasets/umd-zhou-lab/TSRBench |
| Task | System | O Acc. | AO Acc. | C W | W C | Changed |
| Anomaly detection | TimeOmni-1 | 60.70 | 45.61 | 20.64 | 4.98 | 25.62 |
| ChatTS | 62.46 | 54.39 | 8.42 | 0.35 | 8.77 | |
| Time-MQA | 64.91 | 45.61 | 20.73 | 0.36 | 21.09 | |
| TimeOmni-VL | 70.88 | 45.61 | 38.64 | 9.85 | 48.48 | |
| Pattern recognition | TimeOmni-1 | 60.51 | 52.81 | 24.11 | 16.34 | 48.22 |
| ChatTS | 59.91 | 59.23 | 14.13 | 12.84 | 30.50 |
| System | Original-target Acc. | Recomputed Acc. | Correct update | Stale | Other |
| TimeOmni-1 | 41.55 | 72.80 | 78.90 | 14.98 | 6.11 |
| ChatTS | 30.49 | 68.01 | 82.87 | 6.12 | 11.01 |
| Time-MQA | 32.88 | 48.88 | 50.76 | 18.04 | 31.19 |
| TimeOmni-VL | 39.01 | 74.44 | 77.98 | 5.50 | 16.51 |
| System | Intervention | Target changed | Unchanged-target Acc. | Correct update | Stale | Other |
| TimeOmni-1 | Shuffle | 6.7 | 62.40 | 88.89 | 0.00 | 11.11 |
| Flat | 76.1 | 90.62 | 94.12 | 4.90 | 0.98 | |
| Noise | 13.4 | 73.28 | 61.11 | 27.78 | 11.11 | |
| Swap | 50.4 | 53.03 | 62.69 | 23.88 | 13.43 | |
| Zero | 97.8 | 66.67 | 77.10 | 17.56 | 5.34 | |
| ChatTS | Shuffle | 6.7 | 51.20 | 44.44 | 11.11 | 44.45 |
| Level | Criterion | Agreement (%) | Cohen’s | |
| Claim | Factual grounding | 240 | 95.8 | 0.920 |
| Output | Factual grounding | 100 | 98.0 | 0.969 |
| Output | Inference validity | 100 | 89.0 | 0.837 |
| Output | Conclusion consistency | 100 | 94.0 | 0.895 |
| System | Public training signals | Checkpoint constraints |
| TimeOmni-1 | Human-guided CoT-SFT and GRPO, with answer/format rewards and a forecasting-error reward. | Final model available; no released SFT-only or RL-intermediate checkpoint. |
| ChatTS | Alignment and supervised fine-tuning; the training data released with the paper include multi-series, relational, causal, and rationale-like content. | Historical training data and final model available; no released alignment-only checkpoint. |
| Time-MQA | Reported training uses 7,000 TSQA and 3,000 OpenOrca samples; final specialization is a LoRA adapter. | Adapter permits a fixed-base contrast; the original base revision is not pinned and no intermediate adapter is released. |
| TimeOmni-VL | Time-series understanding/generation and CoT-based post-training. | Final model available; no intermediate checkpoints or complete original training-row manifest. |
| Stage | Samples | Multi-series | Cross-series | Causal | Rationale |
| Stage 1 | 105,085 | 66.69 | 6.89 | 2.27 | 51.15 |
| Stage 2 | 30,650 | 29.34 | 19.61 | 10.11 | 42.39 |
| Checkpoint | Original-target Acc. | Recomputed Acc. | Correct update | Stale | Other |
| Base | 39.91 | 27.06 | 17.13 | 43.43 | 39.45 |
| LoRA | 24.81 | 38.27 | 44.04 | 16.51 | 39.45 |
| Checkpoint | Intervention | Eligible | Changed | Original-target Acc. | Recomputed Acc. | Correct update | Stale | Other |
| Base | Shuffle | 134 | 9 | 38.06 | 35.07 | 11.11 | 55.56 | 33.33 |
| Flat | 134 | 102 | 36.57 | 8.96 | 1.96 | 38.24 | 59.80 | |
| Noise | 134 | 18 | 37.31 | 40.30 | 33.33 | 11.11 | 55.56 | |
| Swap | 133 | 67 | 40.60 | 30.08 | 28.36 | 49.25 | 22.39 | |
| Zero | 134 | 131 | 47.01 | 20.90 | 21.37 | 48.09 | 30.54 | |
| LoRA | Shuffle | 134 | 9 | 32.09 | 29.85 | 22.22 | 55.56 | 22.22 |
| Valid (%) | Paired-valid acc. (%) | Overall acc. (%) | ||||
| Task | Base | LoRA | Base | LoRA | Base | LoRA |
| Anomaly detection | 98.95 | 92.63 | 55.17 | 65.90 | 56.84 | 61.05 |
| Pattern recognition | 26.97 | 91.49 | 73.91 | 53.42 | 20.06 | 53.45 |
| Cross-series comparison | 26.69 | 25.73 | 50.00 | 48.68 | 13.73 | 13.35 |
| Causal diagnosis | 50.38 | 86.08 | 23.93 | 35.58 | 11.65 | 25.06 |
| Decision intervention | 36.00 | 90.86 | 24.56 | 24.56 | 9.71 | 21.14 |