cs.DBOct 4, 2026

SALUS: Automated Auditing of NL-to-SQL Benchmarks through Weak Supervision of Multi-Agent Output

Authors: Shiyuan Zhou, Ashwin Gerard Colaco, Sainyam Galhotra, Sharad Mehrotra

Organizations: University of California, Irvine, Irvine, California, USA · Cornell University, Ithaca, New York, USA

Abstract

Natural language to SQL (NL-to-SQL) benchmarks are foundational to progress in data analysis research, yet recent work has shown that widely-used benchmarks contain significant annotation errors. These errors silently corrupt evaluation metrics, penalize correct model output, and distort the field's understanding of state-of-the-art performance. We present SALUS, a system that automatically detects annotation errors in NL-to-SQL benchmarks. SALUS frames benchmark auditing as a weakly supervised error detection: SQL generated by multiple LLM agents drive a suite of complementary weak-labeling functions. By passing this noisy vote matrix through a generative label model, we extract high-confidence training samples without requiring human ground truth. These samples train a decision plane that maps gold SQL query features to per-agent trustworthiness, allowing SALUS to intelligently fuse reliability estimates with raw verdicts for rigorous benchmark error detection. We evaluate on BIRD-Clean-xs, a benchmark of 298 BIRD development tasks with manually verified correctness labels. SALUS achieves F1 = 0.9194, significantly outperforming the state-of-the-art baselines. Applying SALUS to the full development sets, we estimate annotation error rates of approximately 37% on BIRD and 27% on Spider.

Figures & tables

Explore similar work

CardsList
  1. AgentNLQ: A General-Purpose Agent for Natural Language to SQL

    May 18, 2026Olena Bogdanov, Yeunji Jung, Chandra Dhir +5Text-To-SqlSql

  2. SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks

    Apr 20, 2026Mohammadtaher Safarzadeh, Hitesh Laxmichand Patel, Afshin Orojlooyjadid +2Text-To-SqlSyntactic Structure

  3. SOMA-SQL: Resolving Multi-Source Ambiguity in NL-to-SQL via Synthetic Log and Execution Probing

    Jun 9, 2026Sai Ashish Somayajula, Marianne Menglin Liu, Chuan Lei +9Text-To-SqlSql