TQTS-Bench: A Multi-Syntax Benchmark for Text-to-Query over Time-Series Databases
Organizations: College of Computer Science and Electronic Engineering, Hunan University
Abstract
Large language models (LLMs) have significantly advanced natural language querying over relational databases, yet their ability to query time-series databases (TSDBs) remains largely unassessed. Existing benchmarks fail to adequately capture the non-unified query syntaxes, diverse application domains, and unique time-specific query intents inherent to TSDBs. To address this gap, we introduce TQTS-BENCH, a multi-syntax benchmark for evaluating text-to-query capabilities over TSDBs. TQTS-BENCH contains 6,125 high-quality question-answering (QA) pairs spanning 97 TSDBs, 23 distinct query syntaxes, 22 application domains, and 4 types of time-specific query intents. It is constructed through a human-centric AI-assisted workflow, where all QA pairs are carefully reviewed and revised by domain experts to ensure quality and correctness. Extensive evaluations of advanced LLMs and state-of-the-art text-to-query methods reveal challenges in querying TSDBs. Even the best-performing model evaluated, Claude-Opus-5, achieves only 48.98% execution accuracy, while humans reach 87.34%. Error analysis reveals that this performance gap mainly stems from the heterogeneous query syntaxes across different TSDBs, misinterpretation of time-specific intents, and incorrect schema linking. These findings highlight new opportunities to narrow the gap between current LLM capabilities and the requirements of TSDB queries in real-world applications. The benchmark is available at: https://anonymous.4open.science/r/TQTS-Bench-00CD.
Figures & tables
| Method | EX( ) | |||
| Easy | Medium | Hard | Overall | |
| Human performance | *95.71% | *85.52% | *81.94% | *87.34% |
| Open-Source Models | ||||
| Qwen3.8-Flash | 43.75% | 20.07% | 17.79% | 27.04% |
| DeepSeek-V4-Pro | 39.89% | 15.87% | 11.79% | 22.51% |
| GLM-5.3-Flash | 54.30% | 28.88% | 19.15% | 34.61% |
| Domain | EX( ) |
| Targeted | 3.03% |
| Others | 0.00% |
| Setting | Syntax | EX( ) |
| Original | Diverse | 0.00% |
| w/ conversion | SQL | 32.21% |
| BIRD | SQL | 56.91% |
| Spider 2.0-lite | SQL | 22.94% |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| TSDB management system | Representative functions and operators | # TSDB | |
| Window aggregation & resampling (I1) | Temporal change analysis (I2) | ||
| InfluxDB OSS v2 | aggregateWindow() | derivative() | 17 |
| InfluxDB 3 Core | date_bin_gapfill() | / | |
| Prometheus | avg_over_time() | resets() | 13 |
| TimescaleDB | time_bucket_gapfill() | delta() | 8 |
| DolphinDB | resample() | ratios() | 5 |
| Query intent | Definition | Example |
| Time-specific | ||
| I1 : Window aggregation & resampling | Resample or aggregate a time series over time windows or a new sampling grid, yielding results for each window or grid. | “What was Hankyung’s average temperature forecast for each day of June 2017?” |
| I2 : Temporal change analysis | Analyzing temporal changes in values within a single time series. | “What was the largest one-second increase in whole-home active power?” |
| I3 : Time localization | Locate the time point(s) or interval at which an event or state occurs. | “ When was the earliest weather observation in the database?” |
| I4 : Relationship analysis | Analyzing relationships between two or more independently identifiable time series. | “On September 12, at which minutes did Apple’s trading volume exceed Amazon’s ?” |
| Time-agnostic | ||
| Error Type | Subcategory | Representative Example(s) |
| Query syntax errors | Query structure errors | Refer to Fig. 6 (a) and Fig. 18 (a) |
| Function/keyword usage errors | Refer to Fig. 6 (b) and Fig. 18 (b) | |
| Intent understanding errors | I1 : Window aggregation & resampling | Refer to Fig. 7 (b) |
| I2 : Temporal change analysis | Refer to Fig. 19 (a) | |
| I3 : Time localization | Refer to Fig. 19 (b) | |
| I4 : Relationship analysis | Refer to Fig. 19 (c) |