In recent years, time-series question answering (QA) systems have made significant progress. However, generating a correct answer does not show whether retaining the supplied numerical series improves task performance, nor whether the prediction is sensitive to changes in that input. While some systems provide rationales, answer accuracy also does not show whether their numerical claims are grounded in the supplied series or whether the stated inference is valid. In this work, we focus on evaluating four time-series QA systems: TimeOmni-1, ChatTS, TimeOmni-VL, and Time-MQA. First, for three systems with released evaluation data, we reproduce their reported results and compare the performance of the systems with their backbones. Then, we introduce a benchmark named COMMON-TSQA, which collects public evaluation datasets from existing time-series benchmarks and unifies their sample representation, task definitions, and answer schemas, while evaluating each system through its own interface under common evaluation criteria. The evaluation uses the original condition and six interventions while keeping the question and target fixed. Our analysis shows that aggregate performance alone can obscure how systems use numerical evidence. Similar task-level scores can arise despite substantial changes in individual predictions. Some interventions induce simple fallback behavior rather than preserved task ability. We also evaluate rationales for factual grounding, inference validity, and consistency with the final answer. We find that rationales often contain time-series claims unsupported by the input. Moreover, the rationale audit shows that agreement between a rationale and its final answer can coexist with incorrect numerical descriptions or invalid intermediate inferences.
Figures & tables
Figure 1: Final-answer correctness and rationales leave distinct ambiguities. (a) The same answer can be compatible with temporal structure, coarse numerical cues, or non-series cues. (b) Even when a rationale is visible, failures can arise in factual grounding, inference validity, or consistency with the final answer.
Figure 2: The original input and six interventions to the numerical evidence. Non-series input content is fixed across conditions.
Figure 3: Systems and matched backbones on native evaluations. Bars are normalized separately within each task so that the better result equals 100; bar heights are intended only for within-task comparison. Labels report the original metric values.
Task
Metric
Model
Original
Shuffle
Flat
Noise
Swap
Answer-only
Zero
Forecasting
MASE ↓
TimeOmni-1
1.739
1.929
1.321
2.060
26.848
17.418
7.235
ChatTS
–
–
–
–
–
–
–
Time-MQA
2.258
2.287
2.106
2.260
23.380
38.241
12.706
TimeOmni-VL
2.650
3.001
2.604
3.291
36.109
38.546
13.221
Anomaly detection
Accuracy (%) ↑
TimeOmni-1
60.70
55.44
45.61
46.32
45.96
45.61
45.61
ChatTS
62.46
68.88
45.61
58.60
54.39
54.39
45.61
Table 1: Scores of four time-series QA systems under original and six interventions. All conditions are scored against the original ground-truth answer. Bold indicates the best score. “–” indicates that the system produced no scorable forecasting outputs.
System
C → W
W → C
Changed
(%)
(%)
(%)
TimeOmni-1
24.11
16.34
48.22
ChatTS
14.13
12.84
30.50
Time-MQA
17.62
11.88
34.65
TimeOmni-VL
19.69
20.38
47.95
Table 2: Pattern-recognition transitions from Original to Answer-only.
Figure 4: Paired diagnostics of intervention responses. ( a ) Each point is a forecasting sample scorable under both Original and Flat. The horizontal axis is ΔD=DFlat−DOriginal , the change in scaled distance to the fixed Original-history-mean forecast; the vertical axis is ΔMASE=MASEFlat−MASEOriginal . Improve + closer denotes ΔD<0 and ΔMASE<0 ; Degraded + closer denotes ΔD<0 and ΔMASE>0 . ( b ) On cross-series sample–intervention pairs whose correct answer changes, the bars show correct updates and stale answers. Correct updates match the recomputed target, and stale answers match the original ground-truth answer.
Figure 5: Grounded claim rate across six tasks under Original. Dashed lines indicate the overall rates, 49.4% for TimeOmni-1, and 50.1% for TimeOmni-VL.
Factual grounding
Inference validity
Conclusion consistency
TimeOmni-1
2.1
27.3
64.0
TimeOmni-VL
11.2
34.3
92.2
Table 3: Output-level rationale audit under original setting. Values are pass rates (%) for the reported factual-grounding, inference-validity, and conclusion-consistency criteria. Conclusion consistency concerns agreement with the system’s own final answer.
Figure 6: (a) Correct updates and stale answers on 327 cross-series pairs whose targets change after intervention; other errors are reported in Appendix G.3 . (b) Base → LoRA changes in output validity and paired-valid accuracy across discrete tasks.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
System
Backbone
TS input
Training
Evaluation data
TimeOmni-1
Qwen2.5-7B
Text
CoT-SFT → GRPO
TSR-Suite
ChatTS
Qwen2.5-14B
TS encoder
Alignment → SFT
Eval A/B
TimeOmni-VL
BAGEL-7B
TS image
CoT post-training
TSR-Suite (OOD)
Time-MQA
Qwen2.5-7B
Text + context
Continual PT (LoRA)
TSQA
Appendix
Table 4: Systems and original evaluation settings.
System
Split
Task/metric
Reproduced
Paper
Difference
TimeOmni-1
ID
Scenario ACC
88.50
90.70
−2.20
ID
Causality ACC
67.60
69.30
−1.70
ID
Forecasting MAE
15.95
14.30
+1.65
ID
Decision ACC
47.90
47.90
0.00
OOD
Scenario ACC
84.80
87.70
−2.90
OOD
Causality ACC
62.30
64.00
−1.70
Appendix
Table 5: Reproduction results. For MAE, lower is better; all other metrics are higher-is-better.
System / backbone
Split
Task / metric
System
Backbone
Improvement
TimeOmni-1 / Qwen2.5-7B
ID
Scenario ACC
88.50
41.00
+47.50
Causality ACC
67.63
18.75
+48.88
Forecasting MAE
15.95
23.27
+7.32
Decision ACC
47.87
23.94
+23.94
OOD
Scenario ACC
84.76
43.05
+41.71
Causality ACC
62.25
21.13
+41.12
Appendix
Table 6: Performance of time-series QA systems and their matched backbones. Accuracy and score metrics are reported as percentages, forecasting MAE retains its original scale.
Figure 7: Tasks and domains represented in Common-TSQA .
Task
Data sources
Samples n
Domains
Forecasting
GIFT-Eval; SciTS; TSR-Suite (OOD); TIME
605
Energy; environment; finance; healthcare; information technology; transportation; climate
Anomaly detection
TSB-AD-U
285
Environment; healthcare; industry; information technology; synthetic
Pattern recognition
TSAQA; TSRBench
623
Energy; environment; finance; healthcare; information technology; synthetic; transportation
Cross-series comparison
TSAQA; CaTS-Bench
517
Energy; environment; finance; healthcare; information technology; transportation; climate
Causal diagnosis
FactoryBench; SenTSR-Bench; TSR-Suite (OOD)
395
Industry; hydrology
Decision intervention
SenTSR-Bench; TSR-Suite (OOD)
175
Industry; energy
Appendix
Table 7: Tasks, source datasets, and domains represented in Common-TSQA after removing all Open-Meteo weather, CAMS air-quality, and ECB exchange-rate samples.
Table 9: Original–Answer-only transitions across the five discrete tasks. C → W, W → C, and Changed use the joint-valid subset as denominator; all entries are percentages.
Figure 9: Original–Flat item-level forecasting diagnostics for Time-MQA and TimeOmni-VL with available forecasting results. Panels are enlarged relative to the main-text diagnostic for readability. Negative horizontal values indicate movement toward the fixed history-mean forecast; negative vertical values indicate lower MASE under Flat.
System
Original-target Acc.
Recomputed Acc.
Correct update
Stale
Other
TimeOmni-1
41.55
72.80
78.90
14.98
6.11
ChatTS
30.49
68.01
82.87
6.12
11.01
Time-MQA
32.88
48.88
50.76
18.04
31.19
TimeOmni-VL
39.01
74.44
77.98
5.50
16.51
Appendix
Table 10: Pooled target-aware results on the exact-oracle subset. Original-target and recomputed accuracies use all eligible pairs; the remaining columns use target-changed pairs as denominator. All entries are percentages.
System
Intervention
Target changed
Unchanged-target Acc.
Correct update
Stale
Other
TimeOmni-1
Shuffle
6.7
62.40
88.89
0.00
11.11
Flat
76.1
90.62
94.12
4.90
0.98
Noise
13.4
73.28
61.11
27.78
11.11
Swap
50.4
53.03
62.69
23.88
13.43
Zero
97.8
66.67
77.10
17.56
5.34
ChatTS
Shuffle
6.7
51.20
44.44
11.11
44.45
Appendix
Table 11: Target-aware decomposition by system and intervention. Target changed is the share of eligible pairs whose recomputed target changes. Correct-update, stale-answer, and other-error percentages use all target-changed outputs.
Level
Criterion
N
Agreement (%)
Cohen’s κ
Claim
Factual grounding
240
95.8
0.920
Output
Factual grounding
100
98.0
0.969
Output
Inference validity
100
89.0
0.837
Output
Conclusion consistency
100
94.0
0.895
Appendix
Table 12: Blinded human validation of the LLM judge. Agreement compares the LLM-generated labels with independent human annotations on a frozen validation subset covering both systems and all six tasks.
System
Public training signals
Checkpoint constraints
TimeOmni-1
Human-guided CoT-SFT and GRPO, with answer/format rewards and a forecasting-error reward.
Final model available; no released SFT-only or RL-intermediate checkpoint.
ChatTS
Alignment and supervised fine-tuning; the training data released with the paper include multi-series, relational, causal, and rationale-like content.
Historical training data and final model available; no released alignment-only checkpoint.
Time-MQA
Reported training uses 7,000 TSQA and 3,000 OpenOrca samples; final specialization is a LoRA adapter.
Adapter permits a fixed-base contrast; the original base revision is not pinned and no intermediate adapter is released.
TimeOmni-VL
Time-series understanding/generation and CoT-based post-training.
Final model available; no intermediate checkpoints or complete original training-row manifest.
Appendix
Table 13: Public training signals and checkpoint constraints relevant to the discussion. Missing public artifacts limit the comparisons supported by the released systems.
Stage
Samples
Multi-series
Cross-series
Causal
Rationale
Stage 1
105,085
66.69
6.89
2.27
51.15
Stage 2
30,650
29.34
19.61
10.11
42.39
Appendix
Table 14: Selected candidate tags in the ChatTS corpus. Samples gives the stage size; all other entries are percentages within that stage. Rule-based tags overlap and are not gold annotations.
Checkpoint
Original-target Acc.
Recomputed Acc.
Correct update
Stale
Other
Base
39.91
27.06
17.13
43.43
39.45
LoRA
24.81
38.27
44.04
16.51
39.45
Appendix
Table 15: Pooled Time-MQA target-aware Base–LoRA contrast. Original-target and recomputed-target accuracies use all 669 eligible sample–intervention pairs. Correct update, stale, and other use the 327 pairs whose correct target changes. All entries are percentages.
Checkpoint
Intervention
Eligible
Changed
Original-target Acc.
Recomputed Acc.
Correct update
Stale
Other
Base
Shuffle
134
9
38.06
35.07
11.11
55.56
33.33
Flat
134
102
36.57
8.96
1.96
38.24
59.80
Noise
134
18
37.31
40.30
33.33
11.11
55.56
Swap
133
67
40.60
30.08
28.36
49.25
22.39
Zero
134
131
47.01
20.90
21.37
48.09
30.54
LoRA
Shuffle
134
9
32.09
29.85
22.22
55.56
22.22
Appendix
Table 16: Intervention-level Time-MQA target-aware behavior. “Changed” is the number of eligible pairs for which the intervention changes the recomputed target. Correct-update, stale, and other percentages use the changed-target subset.
Valid (%)
Paired-valid acc. (%)
Overall acc. (%)
Task
Base
LoRA
Base
LoRA
Base
LoRA
Anomaly detection
98.95
92.63
55.17
65.90
56.84
61.05
Pattern recognition
26.97
91.49
73.91
53.42
20.06
53.45
Cross-series comparison
26.69
25.73
50.00
48.68
13.73
13.35
Causal diagnosis
50.38
86.08
23.93
35.58
11.65
25.06
Decision intervention
36.00
90.86
24.56
24.56
9.71
21.14
Appendix
Table 17: Time-MQA validity and paired-valid accuracy. Paired-valid accuracy uses only samples for which both checkpoints produce valid semantic outputs.
Time series question-answering (TSQA), in which we ask natural language questions to infer and reason about properties of time series, is a promising yet underexplored capability of foundation models. In this work, we present ARFBench, a TSQA benchmark that evaluates the understanding of multimodal foundation models (FMs) on time series anomalies prevalent in software incident data. ARFBench consists of 750 questions across 142 time series and 5.38M data points from 63 production incidents sourced exclusively from internal telemetry at Datadog. We evaluate leading proprietary and open-source LLMs, VLMs, and time series FMs and observe that frontier VLMs perform markedly better than existing baselines; the leading model (GPT-5) achieves a 62.7% accuracy and 51.9% F1. We next demonstrate the promise of specialized multimodal approaches. We develop a novel TSFM + VLM hybrid prototype which we post-train on a small set of synthetic and real data that yields comparable overall F1 and accuracy with frontier models. Lastly, we find models and human domain experts exhibit complementary strengths. We define a model-expert oracle, a best-of-2 oracle selector over model and expert answers, yielding 82.8% F1 and 87.2% accuracy and establishing a new superhuman frontier for future TSQA models. The benchmark is available at https://huggingface.co/datasets/Datadog/ARFBench.
Stephan Xie, Ben Cohen, Mononito Goswami +6
Machine Learning Department, Carnegie Mellon University, Pittsburgh, PA, USA · Datadog AI Research, New York, NY, USA · Amazon Web Services, Seattle, WA, USA
Time series data in real-world deployments is overwhelmingly irregular. Observations are asynchronous, missing values are informative rather than random, and sampling frequencies vary across sensors and operational windows. However, existing Time Series Question Answering (TSQA) benchmarks mostly assume regularly sampled inputs, leaving a fundamental gap in understanding how large language models (LLMs) and AI agents perform under irregular conditions. To bridge this gap, we introduce IRTS-ToolBench, a benchmark of 1,700 questions spanning 10 task types across 13 domains. IRTS-ToolBench is designed to be used independently by any researcher working on LLM-based irregular time series analysis, providing standardized inputs and a reproducible evaluation protocol. Code can be found in https://github.com/SanhornC/IRTS-ToolBench.
Large language models (LLMs) and time-series language models (TSLMs) are increasingly applied to time-series question answering (TSQA). Unlike text-only QA, TSQA requires models to ground answers in temporal signals whose patterns may occur at different scales, specific time locations, or across separated intervals. However, existing benchmarks are typically organized by task types or high-level reasoning categories, making it difficult to diagnose the underlying signal-level capabilities driving model performance. We introduce TS-Skill, a controlled benchmark for evaluating three composable analytical skills in TSQA: temporal scale selection (SK1), temporal localization (SK2), and cross-interval integration (SK3). TS-Skill provides timestamp-aware questions, broad domain coverage, and human-validated QA quality. To construct the benchmark at scale, we develop SKEvol, a skill-guided agentic framework that combines domain-aware time-series seed generation, skill-controlled question generation, metadata- and code-assisted answer construction, multi-phase signal-grounded verification, and human-in-the-loop curation. Experiments on ten state-of-the-art LLMs and TSLMs reveal substantial and uneven capability gaps across SK1-SK3. In particular, SK3 remains consistently challenging for non-agent models, whereas tool-augmented agents show a selective advantage on standalone SK3. These findings demonstrate that skill-level evaluation can uncover temporal reasoning failures that are obscured by aggregate TSQA scores.
Liying Han, Kang Yang, Oliver Wang +9
University of California, Los Angeles · Samsung Research America · Carnegie Mellon University +2