Natural language to SQL (NL-to-SQL) benchmarks are foundational to progress in data analysis research, yet recent work has shown that widely-used benchmarks contain significant annotation errors. These errors silently corrupt evaluation metrics, penalize correct model output, and distort the field's understanding of state-of-the-art performance. We present SALUS, a system that automatically detects annotation errors in NL-to-SQL benchmarks. SALUS frames benchmark auditing as a weakly supervised error detection: SQL generated by multiple LLM agents drive a suite of complementary weak-labeling functions. By passing this noisy vote matrix through a generative label model, we extract high-confidence training samples without requiring human ground truth. These samples train a decision plane that maps gold SQL query features to per-agent trustworthiness, allowing SALUS to intelligently fuse reliability estimates with raw verdicts for rigorous benchmark error detection. We evaluate on BIRD-Clean-xs, a benchmark of 298 BIRD development tasks with manually verified correctness labels. SALUS achieves F1 = 0.9194, significantly outperforming the state-of-the-art baselines. Applying SALUS to the full development sets, we estimate annotation error rates of approximately 37% on BIRD and 27% on Spider.
Figures & tables
Figure 1 . Probability that the rank-1 system on the BIRD development set truly outperforms lower-ranked methods under the estimated 37% annotation error rate. The rank-1 system cannot be reliably distinguished from the next 20. The plot compares the estimated probability that the rank-1 BIRD development-set system outperforms each lower-ranked system with a 95 percent threshold. Under the estimated 37 percent annotation error rate, comparisons with the next 20 systems do not reach that threshold.
Symbol
Description
B,N,i
Benchmark of N tasks, indexed by i
(qi,si,Di)
NL question, gold SQL, and database for task i
R(s,D)
Result set from executing s on database D
A,K,j,Aj
Set of K LLM agents, indexed by j
s^i,j,vi,j
Candidate SQL and verdict of Aj on task i
mi,j
Validity mask of Aj on task i
Table 1 . Summary of notation.
Figure 2 . An overview of the SALUS pipeline. Benchmark tasks are sent to multiple LLM agents. The question, candidate SQL, execution results, and agent verdicts supply labeling functions whose vote matrix is aggregated by a generative model. High-confidence samples train per-agent reliability models. SQL structural features and agent verdicts enter the decision plane ensemble, which predicts benchmark annotation errors.
Figure 3 . Conceptual demonstration of the comparison of majority vote and the decision plane. Each query (Q1–Q4) shows four agent verdicts where ✓ = agent execution matches gold, ✗ = agent execution differs from gold). Left: majority vote weights all agents equally and predicts Incorrect only when more than half of valid agents disagree with the gold SQL. Right: the decision plane assigns each query to a region of the SQL feature space where a specific agent is most reliable (indicated by the region label and background color), and follows that agent’s verdict. Two panels compare predictions for four queries over join count and aggregation count. Majority voting predicts Q1, Q2, and Q3 correct and Q4 incorrect. Feature-conditioned routing divides this space into four regions assigned to different agents and predicts Q1 correct and Q2, Q3, and Q4 incorrect.
Figure 4 . End-to-end example of SALUS on a single benchmark task. The orange path shows the training flow: agent verdicts and gold SQL produce labeling function votes, which a generative model aggregates into high-confidence labels that supervise the decision plane. The green path shows inference: SQL structural features map the query to a decision plane region, producing per-agent reliability scores that are combined with agent verdicts to predict the final label. The visualized decision plane is a conceptual 2D projection; the actual model operates over V features. A patient-count query illustrates the pipeline. Three agents disagree with the reference and one agrees. Example labeling functions vote incorrect, correct, or abstain. The label model provides weak supervision. SQL features include where count 1, distinct count 1, no limit, and table count 2. The decision plane assigns agent reliability scores of 0.43, 0.72, 0.13, and 0.02; the ensemble predicts an annotation error.
Error Type
# Errors
Question
Explanation
Incorrect SQL
134
List out the account numbers of female clients who are oldest and has lowest average salary…
Gold SQL appends LIMIT 1 , returning only one account when multiple accounts satisfy the condition.
Query Ambiguity
14
Which active district has the highest average score in Reading?
“Average score” can refer to a single school’s score or the mean across all schools in a district; the gold SQL selects a single school rather than aggregating by district.
Data Inconsistency
7
How many posts were created on 21st July, 2010?
Two tables record creation dates for the same posts but contain different row counts; querying different tables yields different results.
Misinterpreting Domain Semantics
8
Provide the hair colour of the human superhero who is 185 cm tall.
The question assumes a unique answer, but multiple human superheroes are 185 cm tall, so the gold SQL returns more than one result.
Incorrect Evidence
3
State the race and year of race in which Michael Schumacher had his fastest lap.
The provided evidence references “Alex Yoong” instead of Michael Schumacher, introducing a factual error into the annotation context.
Total
166
Table 2 . Representative examples of each annotation error type in BIRD-Clean-xs , with the number of errors per category.
Signal Source
Category
Count
Polarity (B/U)
Execution
Result-set comparison
14
5 / 9
Agent consensus
2
2 / 0
Structural
AST comparison
9
3 / 6
Schema consistency
1
0 / 1
Known error patterns
6
0 / 6
Intent
NL–SQL alignment
11
3 / 8
Table 3 . Breakdown of labeling functions by signal source, category, and label polarity.
(A) BIRD-Clean-xs
(B) 200-task random sample
Method
F1
Precision
Recall
F1
Precision
Recall
Direct methods
Majority Vote
0.8895
0.8596
0.9217
0.7843
0.7059
0.8824
Best Single Agent (Opus 4.6)
0.8616
0.9013
0.8253
0.7612
0.7727
0.7500
Weak-labeling methods (transductive)
SALUS
0.9194
0.9112
0.9277
0.8552
0.8052
0.9118
Table 4 . Detection performance on (A) BIRD-Clean-xs and (B) the 200-task random sample.
Baseline
Δ F1
95% CI
p
GPT-4o
+0.154
[0.095, 0.221]
< 0.001
Gemini 2.5 Pro
+0.153
[0.087, 0.225]
< 0.001
GPT-5
+0.097
[0.042, 0.156]
< 0.001
Opus 4.6
+0.094
[0.025, 0.167]
0.002
Majority Vote
+0.071
[0.035, 0.114]
< 0.001
Weak Labels
+0.053
[0.009, 0.100]
0.009
Table 5 . Paired bootstrap significance test ( b=10,000 )
Error Type
Errors
Detected
Recall
95% CI
Incorrect SQL
134
127
0.9478
[0.895, 0.979]
Query Ambiguity
14
14
1.0000
[0.768, 1.000]
Misinterpreting Domain Semantics
8
6
0.7500
[0.349, 0.968]
Data Inconsistency
7
4
0.5714
[0.184, 0.901]
Incorrect Evidence
3
3
1.0000
[0.292, 1.000]
All errors
166
154
0.9277
[0.877, 0.962]
Table 6 . Per-error-type detection of SALUS on BIRD-Clean-xs , with exact (Clopper–Pearson) 95% binomial CIs on recall.
Difficulty
Errors
Detected
F1
Simple
90
83
0.9274
Moderate
58
54
0.9076
Challenging
18
17
0.9189
Table 7 . Per-difficulty detection of SALUS on BIRD-Clean-xs , using difficulty labels provided by the BIRD team.
System
F1
Precision
Recall
SALUS (Frontier)
0.9194
0.9112
0.9277
SALUS (Legacy)
0.8412
0.8218
0.8614
SAR-Agent (o3)
0.8389
0.8466
0.8313
SQLDriller
0.6456
0.6800
0.6145
Table 8 . Detection performance of SALUS and existing auditing systems on BIRD-Clean-xs .
System
F1
Recall
TP
FN
SALUS (Frontier)
0.9559
0.9155
65
6
SALUS (Legacy)
0.8906
0.8028
57
14
SAR-Agent (o3)
0.8992
0.8169
58
13
SQLDriller
0.7321
0.5775
41
30
Table 9 . Detection performance of SALUS and existing auditing systems on the LIMIT / DISTINCT error set ( n = 71).
Error Type
n
Frontier
Legacy
SAR
SQLD
Incorrect SQL
134
0.95 (127)
0.88 (118)
0.85 (114)
0.63 (84)
Query Ambiguity
14
1.00 (14)
0.86 (12)
0.79 (11)
0.64 (9)
Misinterp. Domain Sem.
8
0.75 (6)
0.75 (6)
0.75 (6)
0.75 (6)
Data Inconsistency
7
0.57 (4)
0.57 (4)
0.71 (5)
0.43 (3)
Incorrect Evidence
3
1.00 (3)
1.00 (3)
0.67 (2)
0.00 (0)
All error types
166
0.93 (154)
0.86 (143)
0.83 (138)
0.61 (102)
Table 10 . Per-error-type recall on BIRD-Clean-xs ; the last row is specificity. Frontier and Legacy denote SALUS configurations; SAR and SQLD denote SAR-Agent and SQLDriller.
System
Calls
Total Tokens
Cost ($)
Latency
Total $
SALUS (Frontier)
4.0
5,499
0.0332
15.3 s
50.89
SALUS (Legacy)
4.0
3,402
0.0039
2.6 s
6.04
SQLDriller (GPT-4.1)
14.6
18,100
0.0449
149.7 s
68.91
SAR-Agent (o3)
7.7
36,966
0.0680
39.9 s
104.38
Table 11 . Cost and latency on BIRD . Values are per task except total cost, projected to all 1,534 tasks from the 50-task sample.
Benchmark
Det. Rate
Est. True Rate
95% CI
BIRD dev (SQLDriller)
42.8% (656/1534)
≈ 39%
[33%, 45%]
BIRD dev (Nov. 2025)
37.2% (571/1534)
≈ 37%
[30%, 43%]
Spider dev
31.4% (325/1034)
≈ 27%
[21%, 33%]
Table 12 . Benchmark-wide error rate estimation. Each row uses TPR and FPR from a manually labeled 200-task sample of the corresponding benchmark version.
Statistic
Rate
Agent vs. gold disagreement
GPT-5 vs. Gold (highest)
45.4%
Opus 4.6 vs. Gold (lowest)
36.7%
Pairwise disagreement between agents
GPT-5 vs. GPT-4o (highest)
37.1%
Opus 4.6 vs. Gemini 2.5 Pro (lowest)
30.4%
Table 13 . Agent–gold and inter-agent disagreement rates on the BIRD development set. Only the highest and lowest rates are shown for each category;
Ablation Configuration
F1
Δ F1
Single-group removals
None (full system)
0.9194
—
− Structural signals
0.8758
− 0.0436
− Execution signals
0.8829
− 0.0365
− Meta LF
0.8976
− 0.0218
− Intent signals
0.9024
− 0.0170
Table 14 . Ablation of LF groups on BIRD-Clean-xs , including every valid multi-group removal.
Agent
Top Features
GPT-5
has_subquery , agg_count
GPT-4o
join_count , table_count
Opus 4.6
join_count , agg_count , has_distinct
Gemini 2.5 Pro
table_count , cardinality
Table 15 . Top features in the depth-2 CART reliability model for each frontier agent, ranked by Gini importance.
Quantity
SALUS
SQLDriller
Errors detected (of 134)
127 (94.8%)
84 (62.7%)
Usable correction proposals
75
69
Execution-match with manual fix
59
35
Conditional accuracy
78.7%
50.7%
End-to-end correct fix rate
44.0%
26.1%
Table 16 . Automatic correction of the 134 Incorrect SQL errors; correctness requires execution match with the manual correction.
Natural language to SQL (NL2SQL) conversion is an important problem for researchers and enterprises due to the ubiquitous importance of relational databases in broad-ranging practical problems. Despite the rapid advancements in the capabilities of LLMs, NL2SQL has not reached parity in accuracy with human expert SQL writers, hence needing additional improvements in NL2SQL algorithms. This study presents a new multi-agent method for NL2SQL that achieves 78.1% semantic accuracy on the BIg Bench for LaRge-scale Database (BIRD) benchmark. Our method leverages a semantically enriched representation of user-provided schema, adds user-provided business rules, and produces accurate SQL queries. The main contributions of this study are (a) We designed an optimized new orchestrator in a multi-agent solution that uses LLMs to plan, orchestrate, reflect, and self-correct to generate accurate SQL queries, (b) We developed an advanced schema enrichment method that creates context-aware metadata to improve accuracy, and (c) We demonstrated the accuracy and generalizability of the method across different domains and datasets by evaluating it on the BIRD-SQL benchmark.
Large language models (LLMs) have achieved strong performance on natural language to SQL (NL2SQL) benchmarks, yet their reported accuracy may be inflated by contamination from benchmark queries or structurally similar patterns seen during training. We introduce SPENCE (Syntactic Probing and Evaluation of NL2SQL Contamination Effects), a controlled syntactic probing framework for detecting and quantifying such contamination. SPENCE systematically generates syntactic variants of test queries for four widely used NL2SQL datasets-Spider, SParC, CoSQL, and the newer BIRD benchmark. We use SPENCE to evaluate multiple high-capacity LLMs under execution-based scoring. For each model, we measure changes in execution accuracy across increasing levels of syntactic divergence and quantify rank sensitivity using Kendall's tau with bootstrap confidence intervals. By aligning these robustness trends with benchmark release dates, we observe a clear temporal gradient: older benchmarks such as Spider exhibit the strongest negative values and thus the highest likelihood of training leakage, whereas the more recent BIRD dataset shows minimal sensitivity and appears largely uncontaminated. Together, these findings highlight the importance of temporally contextualized, syntactic-probing evaluation for trustworthy NL2SQL benchmarking.
Natural language interfaces to databases aim to translate user questions into executable SQL, yet remain brittle in real-world settings where questions are underspecified and schemas are large and ambiguous. Ambiguity across user questions, database schemas, and model interpretations are central failure modes in NL2SQL, leading to misaligned intent, incorrect schema grounding, and erroneous SQL generation. Existing approaches rely on human clarification or treat ambiguity as a schema representation problem, but these do not scale nor resolve ambiguity autonomously. We propose SOMA-SQL to automatically resolve ambiguity via targeted synthetic query log and ambiguity-driven probing. SOMA-SQL constructs synthetic query log to ground schema interpretation and guide candidate SQL generation; it then executes targeted probing queries, driven by a structured ambiguity taxonomy and candidate disagreements, to produce disambiguation evidence for final SQL selection and repair. This active approach to ambiguity discovery and resolution generalizes across unseen schemas and query distributions without human-in-the-loop. Experiments on six public benchmarks demonstrate that SOMA-SQL improves execution accuracy by 13.0% on average over state-of-the-art baselines, with gains of up to 16.7% on ambiguous questions.
Sai Ashish Somayajula, Marianne Menglin Liu, Chuan Lei +9