Large language models (LLMs) have significantly advanced natural language querying over relational databases, yet their ability to query time-series databases (TSDBs) remains largely unassessed. Existing benchmarks fail to adequately capture the non-unified query syntaxes, diverse application domains, and unique time-specific query intents inherent to TSDBs. To address this gap, we introduce TQTS-BENCH, a multi-syntax benchmark for evaluating text-to-query capabilities over TSDBs. TQTS-BENCH contains 6,125 high-quality question-answering (QA) pairs spanning 97 TSDBs, 23 distinct query syntaxes, 22 application domains, and 4 types of time-specific query intents. It is constructed through a human-centric AI-assisted workflow, where all QA pairs are carefully reviewed and revised by domain experts to ensure quality and correctness. Extensive evaluations of advanced LLMs and state-of-the-art text-to-query methods reveal challenges in querying TSDBs. Even the best-performing model evaluated, Claude-Opus-5, achieves only 48.98% execution accuracy, while humans reach 87.34%. Error analysis reveals that this performance gap mainly stems from the heterogeneous query syntaxes across different TSDBs, misinterpretation of time-specific intents, and incorrect schema linking. These findings highlight new opportunities to narrow the gap between current LLM capabilities and the requirements of TSDB queries in real-world applications. The benchmark is available at: https://anonymous.4open.science/r/TQTS-Bench-00CD.
Figures & tables
Figure 1: The three key differences between RDBs and TSDBs: (a) Syntaxes : unified SQL vs. diverse, non-uniform query syntaxes. (b) Domains : narrow domain coverage vs. broad domain coverage. (c) Intents : time-agnostic operations vs. time-specific analysis.
Figure 2: The construction pipeline of TQTS-Bench .
Figure 3: A representative question expressed using Flux, PromQL, and QuestDB SQL.
Figure 4: (a) Domain coverage and (b) time-specific query intent statistics of TQTS-Bench .
Method
EX( ↑ )
Easy
Medium
Hard
Overall
Human performance
*95.71%
*85.52%
*81.94%
*87.34%
Open-Source Models
Qwen3.8-Flash
43.75%
20.07%
17.79%
27.04%
DeepSeek-V4-Pro
39.89%
15.87%
11.79%
22.51%
GLM-5.3-Flash
54.30%
28.88%
19.15%
34.61%
Table 1: Evaluation results on TQTS-Bench . The best result (except human) is in bold , and the runner-up is underlined . Results with (*) are tested on a randomly sampled 10% subset.
Domain
EX( ↑ )
Targeted
3.03%
Others
0.00%
Table 2: Performance of PromCopilot on different domains.
Figure 5: Proportions of incorrect cases by error type.
Figure 6: Examples of query syntax errors . (a) Query structure error: reversing group and aggregateWindow selects per-series rather than per-group last readings, producing incorrect results. (b) Function usage error: replacing bar(Time, 15m) with floor(Time, 15m) violates the single-argument requirement of floor , causing a compilation error.
Figure 7: Intent understanding errors . (a) Distribution of primary erroneous intents; time-specific intents account for 65.41%. (b) Window aggregation & resampling error: the predicted query lacks [2h:15m] , returning a single top-3 result instead of results at each 15-minute checkpoint.
Setting
Syntax
EX( ↑ )
Original
Diverse
0.00%
w/ conversion
SQL
32.21%
BIRD
SQL
56.91%
Spider 2.0-lite
SQL
22.94%
Table 3: EX across different query syntaxes and benchmarks.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
TSDB management system
Representative functions and operators
# TSDB
Window aggregation & resampling (I1)
Temporal change analysis (I2)
InfluxDB OSS v2
aggregateWindow()
derivative()
17
InfluxDB 3 Core
date_bin_gapfill()
/
Prometheus
avg_over_time()
resets()
13
TimescaleDB
time_bucket_gapfill()
delta()
8
DolphinDB
resample()
ratios()
5
Appendix
Table 4: Selected TSDB management systems and their temporal functions and operators.
Figure 8: The domain and size of the TSDBs in TQTS-Bench .
Query intent
Definition
Example
Time-specific
I1 : Window aggregation & resampling
Resample or aggregate a time series over time windows or a new sampling grid, yielding results for each window or grid.
“What was Hankyung’s average temperature forecast for each day of June 2017?”
I2 : Temporal change analysis
Analyzing temporal changes in values within a single time series.
“What was the largest one-second increase in whole-home active power?”
I3 : Time localization
Locate the time point(s) or interval at which an event or state occurs.
“ When was the earliest weather observation in the database?”
I4 : Relationship analysis
Analyzing relationships between two or more independently identifiable time series.
“On September 12, at which minutes did Apple’s trading volume exceed Amazon’s ?”
Time-agnostic
Appendix
Table 5: The definitions and examples of the nine query intents used in TQTS-Bench .
Figure 9: The statistics of query intents. (a) Distribution of query intents across different difficulty levels; (b) Distribution of query intents in TQTS-Bench ; (c) Average query intent count across different difficulty levels.
Figure 10: The seven types of schema structures of the same time-series data.
Figure 11: An example of the customized Prometheus database context used in TQTS-Bench .
Figure 12: The visual interface for annotators in human-AI collaboration step.
Figure 13: The visual interface for adjudicator in QA verification step.
Figure 14: The evaluation criteria and an example of the match policy used in the evaluation.
Figure 15: An excerpt from the SQL generation prompt of DeepEye-SQL.
Figure 16: An example where DeepEye-SQL generates SQL instead of the required PromQL.
Figure 17: The description of tightly coupled domains as reflected in (a) the prompt for query generation and (b) the statement of external validity in the original paper.
Error Type
Subcategory
Representative Example(s)
Query syntax errors
Query structure errors
Refer to Fig. 6 (a) and Fig. 18 (a)
Function/keyword usage errors
Refer to Fig. 6 (b) and Fig. 18 (b)
Intent understanding errors
I1 : Window aggregation & resampling
Refer to Fig. 7 (b)
I2 : Temporal change analysis
Refer to Fig. 19 (a)
I3 : Time localization
Refer to Fig. 19 (b)
I4 : Relationship analysis
Refer to Fig. 19 (c)
Appendix
Table 6: Overview of the error types and their representative examples.
Figure 18: Examples of query syntax errors. (a) Query structure error: placing LAG() directly in the WHERE clause, rather than computing the preceding reading in a subquery, causes a window-function execution error. (b) Operator usage error: replacing subtraction with offset 1h by delta(...[1h]) substitutes an extrapolated change for the difference between two time points; together with a changed metric, this produces incorrect values and rankings.
Figure 19: Examples of intent understanding errors. (a) Temporal change analysis error: ranking signed differences rather than absolute magnitudes interprets the largest changes as the largest increases, selecting incorrect time boundaries. (b) Time localization error: returning available-memory values instead of applying max(timestamp(...)) reports memory quantities rather than the latest sample time. (c) Relationship analysis error: using IN (8, 32) instead of matching timestamps across both communities confuses union with intersection; a global LIMIT 1 further returns a single time rather than the earliest shared time for each day.
Figure 20: Two examples of schema linking errors. (a) The model treats the attack_number tag as a field and filters on _field == "attack_number" , which removes all relevant records and produces an empty result. (b) The model uses designated_timestamp instead of the metadata column designatedTimestamp , causing an invalid-column error and preventing retrieval of the designated timestamp column.
Figure 21: A case study on the differences in LLM capabilities between generating Flux for TSDBs and SQL for RDBs.
Large language models (LLMs) and time-series language models (TSLMs) are increasingly applied to time-series question answering (TSQA). Unlike text-only QA, TSQA requires models to ground answers in temporal signals whose patterns may occur at different scales, specific time locations, or across separated intervals. However, existing benchmarks are typically organized by task types or high-level reasoning categories, making it difficult to diagnose the underlying signal-level capabilities driving model performance. We introduce TS-Skill, a controlled benchmark for evaluating three composable analytical skills in TSQA: temporal scale selection (SK1), temporal localization (SK2), and cross-interval integration (SK3). TS-Skill provides timestamp-aware questions, broad domain coverage, and human-validated QA quality. To construct the benchmark at scale, we develop SKEvol, a skill-guided agentic framework that combines domain-aware time-series seed generation, skill-controlled question generation, metadata- and code-assisted answer construction, multi-phase signal-grounded verification, and human-in-the-loop curation. Experiments on ten state-of-the-art LLMs and TSLMs reveal substantial and uneven capability gaps across SK1-SK3. In particular, SK3 remains consistently challenging for non-agent models, whereas tool-augmented agents show a selective advantage on standalone SK3. These findings demonstrate that skill-level evaluation can uncover temporal reasoning failures that are obscured by aggregate TSQA scores.
Liying Han, Kang Yang, Oliver Wang +9
University of California, Los Angeles · Samsung Research America · Carnegie Mellon University +2
Time series data in real-world deployments is overwhelmingly irregular. Observations are asynchronous, missing values are informative rather than random, and sampling frequencies vary across sensors and operational windows. However, existing Time Series Question Answering (TSQA) benchmarks mostly assume regularly sampled inputs, leaving a fundamental gap in understanding how large language models (LLMs) and AI agents perform under irregular conditions. To bridge this gap, we introduce IRTS-ToolBench, a benchmark of 1,700 questions spanning 10 task types across 13 domains. IRTS-ToolBench is designed to be used independently by any researcher working on LLM-based irregular time series analysis, providing standardized inputs and a reproducible evaluation protocol. Code can be found in https://github.com/SanhornC/IRTS-ToolBench.
Large Language Models (LLMs) are transforming database interaction paradigms, evolving from simple query translators to autonomous database administrators (DBAs). However, current evaluation benchmarks remain disproportionately fixated on Text-to-SQL tasks, neglecting the holistic Database Lifecycle-from initial schema design to post-deployment maintenance. This narrow focus fails to capture the diverse capabilities required for real-world database management. To bridge this gap, we introduce DBLifeBench, the first benchmark to evaluate LLMs across five critical lifecycle phases: Design, Implementation, Operation, Debugging, and Maintenance. Furthermore, addressing the cognitive mismatch between ambiguous natural language and complex SQL logic, we propose Progressive-Text2SQL, a novel task utilizing structured reasoning graphs to mimic human iterative problem-solving. Our extensive evaluation reveals a critical insight: while general-purpose models demonstrate balanced performance, specialized Text-to-SQL models suffer from ``catastrophic forgetting'' in non-coding phases like design and maintenance. DBLifeBench serves as a foundational step toward evaluating and building true full-stack database intelligence.
Shunfan Zheng, Dongsheng Shi, Yue Li +3
1East China Normal University · 2Hasso Plattner Institute/University of Potsdam