cs.CLSep 28, 2026

TQTS-Bench: A Multi-Syntax Benchmark for Text-to-Query over Time-Series Databases

Authors: Fei Lyu, Zhiyi Peng, Jiaming Liu, Yixuan Yang, Changjian Chen, Zhuo Tang, Jiapeng Zhang, Kenli Li

Organizations: College of Computer Science and Electronic Engineering, Hunan University

Abstract

Large language models (LLMs) have significantly advanced natural language querying over relational databases, yet their ability to query time-series databases (TSDBs) remains largely unassessed. Existing benchmarks fail to adequately capture the non-unified query syntaxes, diverse application domains, and unique time-specific query intents inherent to TSDBs. To address this gap, we introduce TQTS-BENCH, a multi-syntax benchmark for evaluating text-to-query capabilities over TSDBs. TQTS-BENCH contains 6,125 high-quality question-answering (QA) pairs spanning 97 TSDBs, 23 distinct query syntaxes, 22 application domains, and 4 types of time-specific query intents. It is constructed through a human-centric AI-assisted workflow, where all QA pairs are carefully reviewed and revised by domain experts to ensure quality and correctness. Extensive evaluations of advanced LLMs and state-of-the-art text-to-query methods reveal challenges in querying TSDBs. Even the best-performing model evaluated, Claude-Opus-5, achieves only 48.98% execution accuracy, while humans reach 87.34%. Error analysis reveals that this performance gap mainly stems from the heterogeneous query syntaxes across different TSDBs, misinterpretation of time-specific intents, and incorrect schema linking. These findings highlight new opportunities to narrow the gap between current LLM capabilities and the requirements of TSDB queries in real-world applications. The benchmark is available at: https://anonymous.4open.science/r/TQTS-Bench-00CD.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering

    May 23, 2026Liying Han, Kang Yang, Oliver Wang +9Temporal ReasoningSkills

  2. Towards Verifiable Agentic Data Science: Solving Irregular TSQA Via Tool-Grounded Reasoning

    Jun 13, 2026Sanhorn Chen, Xiaoyang Chen, Boyu Liu +1Time SeriesData Science Agents

  3. Evaluating LLMs in Database Scenarios: A Lifecycle Benchmark for Assessing Their Potential in Core Database Tasks

    Aug 4, 2026Shunfan Zheng, Dongsheng Shi, Yue Li +3Text-To-SqlLarge Language Model Benchmarks