Text-to-SQL systems are typically trained and evaluated on a single dialect (SQLite), yet production deployments span PostgreSQL, MySQL, ClickHouse, and beyond. We show that this single-dialect assumption leads to a substantial drop in cross-dialect accuracy for every model we tested. The drop persists across scale, architecture, and even purpose-built text-to-SQL systems. We argue that the fix is to change the generation target: instead of asking an LLM to emit dialect-specific SQL, we have it emit a dialect-agnostic relational algebra query plan, which a deterministic compiler then renders into SQL for any supported backend. Across thirteen models from 3B to frontier scale, this restores cross-dialect portability nearly uniformly, at a small cost in peak accuracy on the model's home dialect for capable prompted models and none once fine-tuned on plans; under matched fine-tuning, plan supervision yields a stronger model than SQL supervision. We also introduce MetricName, a question-aware result-set comparator needed to evaluate fairly across dialects, where existing metrics confound semantic errors with benign cross-dialect variation. More broadly, the result is a reminder that a generation target chosen for execution is not necessarily the one that maximizes generation quality.
Figures & tables
SQLite
PG
MySQL
CH
BIRD
97.1
95.4
95.6
94.7
Spider
97.2
96.7
96.6
95.6
Table 1: Round-trip conversion accuracy (%) of the pipeline under DF-Match. PG = PostgreSQL, CH = ClickHouse. Each gold SQLite query is converted to a Calcite plan and back to the target dialect; accuracy measures whether the reconstructed query returns an equivalent result as the original.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
BIRD-dev
Spider-dev
Model
Target
SQLite
PG
MySQL
CH
Port.
SQLite
PG
MySQL
CH
Port.
Small
Ministral 3-8B †
SQL
61.2
52.4
57.4
50.3
82.1
85.3
78.3
81.5
77.6
90.9
Ministral 3-8B †
QP
64.5
64.2
65.2
63.5
97.3
84.1
83.2
84.5
82.9
98.2
SQLCoder
SQL
39.1
29.3
35.1
27.6
70.6
66.0
18.4
50.3
18.2
27.5
OmniSQL
SQL
62.1
54.9
58.0
48.7
78.5
86.0
76.1
83.0
76.5
88.5
Appendix
Table 2: Per-dialect execution accuracy (%) on BIRD-dev and Spider-dev under DF-Match; all prompted runs use greedy decoding (single sample). PG = PostgreSQL, CH = ClickHouse. Port. = portability, defined as worst-dialect / best-dialect accuracy (higher = more portable; 100.0 = identical across dialects). † LoRA fine-tuned.
Metric
Count
%
DF-Match
99
96.1
BIRD-EA
0 4
0 3.9
Appendix
Table 3: Human agreement on 103 BIRD cases where DF-Match and BIRD-EA disagree. McNemar χ2=85.79 , p=2×10−20 .
Metric
SQLite
PG
MySQL
CH
BIRD-EA
83.1
81.2
79.7
80.6
DF-Match
97.1
95.4
95.6
94.7
Gap
+14.0
+14.2
+15.9
+14.1
Appendix
Table 4: Round-trip conversion reliability (%) on BIRD under DF-Match vs. BIRD-EA . PG = PostgreSQL, CH = ClickHouse. The consistent 14–16 pp gap per dialect is a metric artefact, not a pipeline defect.
Existing text-to-SQL benchmarks are largely centered on SQLite, making it difficult to evaluate whether models can generalize across heterogeneous SQL dialects. However, real-world database systems differ substantially in syntax, functions, type systems, and execution semantics, so the same natural language intent often requires dialect-specific SQL realizations. We introduce UniQL, a human-verified benchmark for cross-dialect text-to-SQL evaluation. UniQL aligns 1,534 natural language questions with executable SQL annotations across 16 SQL dialects, yielding 24,544 dialect-specific queries. All dialects share the same intents, aligned schemas and database contents, enabling controlled evaluation of dialect generalization. UniQL is constructed through a hybrid pipeline combining database migration, SQL translation, execution-guided verification, iterative rule summarization, and human validation. Experiments on both open-source and closed-source LLMs show that current models remain far from dialect-universal, with substantial performance variation across database systems and limited transfer from SQLite success to other dialects. These findings highlight the need for aligned cross-dialect benchmarks and more dialect-aware text-to-SQL methods. Code and data are available at https://github.com/JerryGao818/UniQL
Jianling Gao, Chongyang Tao, Jiayuan Bai +7
1SKLCCSE, Beihang University · 2The University of Hong Kong
SQL dialects vary in syntax, types, and functions across database engines. Text-to-SQL benchmarks, however, predominantly support only SQLite. This creates a critical evaluation gap: cross-dialect evaluation reveals weak per-query agreement (Cohen's ), showing that SQLite performance is an unreliable proxy for other dialects. Yet such evaluation remains prohibitively difficult: existing approaches either require expensive manual query transpilation or rely on tools that often fail on complex SQL. To close this gap, we introduce PolySQL, a novel dual-execution method that eliminates the need for query transpilation by comparing normalized execution results. Notably, our approach achieves higher evaluation fidelity than query transpilation with 100% query coverage. PolySQL comprises three datasets, enabling the first large-scale cross-dialect study. Our study reveals a 10.1% average accuracy drop from SQLite to other dialects and identifies a significant dialect difficulty hierarchy. We find this degradation stems from logical rather than syntactic errors (61% vs. 8%). We release our framework code and leaderboard to enable rigorous dialect-robust evaluation.
Prompting-based (i.e., non-fine-tuning) Text-to-SQL methods, where underlying large language model parameters are not changed for the task, face three problems: (i) relying on coarse-grained schema information that may not reveal the fine-grained relationships needed to distinguish ambiguous columns, (ii) failing to capture recurring SQL-generation failures, and (iii) suffering from omission or hallucination of components in complex questions. This paper develops DexterSQL, a prompting/non-fine-tuning-based Text-to-SQL system that improves SQL generation with three novel components: (i) deep schema explorator that identifies ambiguous columns, analyzes their individual and joint data distributions to uncover their relationships and the distinct role of each, (ii) database-agnostic rule creator that mines mismatches between generated and gold SQL only on the training database and converts them into database-agnostic corrective rules that capture recurring LLM failure patterns; and (iii) multi-path SQL generation that introduces a dependency-tree-based intermediate representation that uses the question's sentence structure to guide its decomposition into an SQL skeleton for final SQL generation. DexterSQL achieves a higher accuracy compared to the state-of-the-art using both open-source/weight and closed-source/weight models. Particularly, DexterSQL shows a high improvement of at least 5.5% using an open-weight model (GPT-OSS-120B) on BIRDDev, with total accuracy 70.4%. DexterSQL also shows better improvement of at least 1.4% using closed-weight models, with total accuracy 72.1% and 72.9% on BIRD-Dev with GPT-4o and GPT-5.2.
Anik Pramanik, Murat Kantarcioglu, Vincent Oria +1
New Jersey Institute of Technology, USA. · Virginia Tech, USA.