Text-to-SQL models are commonly trained to map questions directly to static queries, whereas real-world database agents operate through stateful, multi-turn interaction with live databases -- inspecting schemas, executing probe queries, diagnosing errors, and revising hypotheses. This creates a critical train-deploy mismatch, as the execution harness that mediates this interaction is introduced only at inference time. To bridge this gap, we propose HarnessSQL, a harness-native post-training framework that preserves the full interaction structure throughout both supervised fine-tuning and reinforcement learning. HarnessSQL builds isolated, executable database environments paired with hidden execution oracles, rolls out teachers directly inside the target SQL harness, and retains only verified trajectories for full-sequence SFT, followed by execution-reward RL. Across Spider 2.0-SQLite, HarnessSQL dramatically boosts the execution accuracy of compact models, raising Qwen3-8B from 15.5% to 45.2% and Qwen3-14B from 22.2% to 54.8%, while transferring effectively to out-of-distribution interactive benchmarks such as BIRD-Interact and LiveSQLBench. Our findings demonstrate that training database agents directly within their execution harness is essential for mastering complex, long-horizon database workflows.
Figures & tables
Figure 1: From one-shot Text-to-SQL to harness-interactive database agents. A stable sandbox and structured SQL harness enable iterative exploration, execution feedback, revision, and submission.
Figure 2: Overview of HarnessSQL , our harness-native post-training framework for interactive SQL agents. (a) Databases and hidden execution oracles are encapsulated into isolated, containerized sandboxes, where a SQL-specific harness mediates agent–environment interaction through four actions: table discovery ( sql_list_tables ), schema inspection ( sql_schema ), query execution ( sql_exec ), and submission ( sql_submit ). The harness exposes schemas, query results, and execution errors while enforcing interaction and execution constraints. Harness-native SFT provides an interaction cold start from verified multi-turn trajectories, after which harness-native RL further optimizes the policy through online sandbox interaction and execution-based rewards. (b) An example trajectory illustrates the resulting workflow: the agent discovers tables, inspects the schema, executes a query, observes an execution error, revises its solution, and submits the verified result.
Method
Model Size
Agentic?
Reasoning?
Harness-trained?
S2-SQLite
BI-mini
LSB-SQLite
General-purpose models
Qwen2.5-Coder
32B
✗
✓
✗
16.3%
3.0%
7.4%
SQL-specialized models
OmniSQL
7B
✗
✗
✗
13.3%
1.7%
7.0%
OmniSQL
32B
✗
✗
✗
14.8%
2.7%
8.9%
Arctic-Text2SQL
7B
✗
✓
✗
15.6%
2.3%
7.0%
Table 1: Execution accuracy across in-domain (S2-SQLite) and out-of-distribution (BI-mini, LSB-SQLite) benchmarks (S2-SQLite: Spider 2.0-SQLite; BI-mini: BIRD-Interact Mini; LSB-SQLite: LiveSQLBench Base-Lite SQLite). Bold indicates the best, and underline indicates the second best.
Training Stage
8B
14B
Base
15.5%
22.2%
Direct RL
20.0%
-
SFT
30.4%
37.8%
SFT + RL
45.2%
54.8%
Table 2: Effect of harness-native post-training on Spider 2.0-SQLite execution accuracy.
Post-training
Inference
S2-SQLite (ID)
BI-mini (OOD)
LSB-SQLite (OOD)
None
Non-harness
2.2%
1.0%
3.3%
None
DSH-SQL
15.5%
3.3%
12.2%
Non-harness
Non-harness
14.1%
1.3%
3.7%
Non-harness
DSH-SQL
9.6%
3.0%
8.2%
Harness-native
DSH-SQL
30.4%
5.3%
16.6%
Table 3: Comparison of non-interactive vs. harness-native training and inference across ID and OOD benchmarks.
Figure 3: Effect of multi-turn trajectory organization for SFT.
Figure 4: Effect of the maximum trajectory context length. The same context limit is used during both SFT and RL.
RL Algorithm
S2-SQLite
BI-mini
LSB-SQLite
GRPO
40.7%
8.7%
23.3%
GSPO
43.7%
9.0%
25.2%
DAPO
45.2%
9.0%
24.8%
Table 4: Comparison of RL algorithms under matched training and rollout budgets.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Layer
Ownership
Responsibility
Visible
Experiment controller
Ours, outside DSH
Loads tasks; creates a process, workspace, and DSH home per task; sets environment bindings; enforces episode deadlines; schedules workers; recovers submissions; records all outcomes.
No
Profile composition
DSH + ours
Composes dsh-base , dsh-headless , and dsh-bundle-sql ; disables unrelated coding-agent capabilities; mounts the selected model and database backend.
No
Control plane
DSH
Prompt and model routing, agent/session state, durable checkpoints, cancellation, retry, compaction, result retention, and headless termination.
Indirect
ReAct-style loop
DSH
Alternates model requests with validated tool dispatch until a terminal response, maximum-token stop, error, or cancellation.
Re-executes the submitted artifact, compares normalized results with the hidden oracle, and checks protocol predicates.
No
Appendix
Table 5: Layers of the evaluation harness. “Visible” indicates whether the model can invoke or directly observe the component, rather than whether its effects can influence the trajectory.
Mechanism
Observed
Interpretation
Session policy initialization
138
Permission preset, sandbox mode, and approval policy were recorded once per session.
sql_list_tables / sql_schema
153 / 199
Both tools appeared in every audited session.
sql_exec / sql_submit
3,230 / 129
All sessions explored with execution; 129 reached explicit submission.
Automatic context compaction
22
The pressure-triggered summarization branch fired only on long trajectories.
Tool-result pruning
1
One over-threshold observation was deterministically shortened.
Repeated-call reminders
19
Middleware injected feedback after repeated identical calls.
Appendix
Table 6: Runtime activation audit for the SQLite profile. A mounted hook may be traversed without its conditional action firing.
Teacher Model
Harness
S2-SQLite
Base Model
Full DSH
13.3%
Qwen3.5-9B
DSH-SQL
7.4%
Qwen3.6-27B
Full DSH
15.5%
Qwen3.6-27B
DSH-SQL
23.0%
Appendix
Table 7: Ablation on trajectory synthesis configurations evaluated on Spider 2.0-SQLite, using Qwen3-8B fine-tuned on 1,000 synthesized trajectories.
Metric
Base
SFT
SFT+DAPO
Trajectories with invalid calls ↓
67.41%
17.41%
0.56%
SQL execution-error rate ↓
17.93%
9.03%
8.08%
Immediate recovery ↑
44.09%
67.80%
72.72%
Final SQL executed before submit ↑
47.78%
70.56%
72.04%
Exactly one valid successful submit ↑
18.70%
75.74%
93.52%
Strict protocol violation ↓
83.15%
26.67%
11.48%
Appendix
Table 8: Interaction behavior across post-training stages on Spider 2.0-SQLite.
Hyperparameter / Setting
Value / Specification
Base models
Qwen3-8B / Qwen3-14B
Training examples
2,512 verified trajectories
Maximum serialized length
32,768 tokens
Loss mask
Assistant reasoning and actions only; system, user, and environment tokens masked
Table 9: SFT configurations used for both Qwen3-8B and Qwen3-14B checkpoints. Both models share identical training settings and are initialized from their respective base checkpoints. Each example is one complete verified harness episode.
Setting
Configuration (Qwen3-8B / Qwen3-14B)
Initialization
32K Qwen3-8B / 14B SFT checkpoint
Rollout horizon
150 rollout iterations
Rollout group
4 prompts per iteration; 8 trajectories per prompt
Updates and batching
2 optimizer steps per rollout; global batch 16; micro-batch 1
Context limits
32,768 tokens per DSH session; at most 4,096 new tokens per model call
Table 10: DAPO configurations used for the reported Qwen3-8B and Qwen3-14B results. “Prompt pool” is the number of records in the corresponding launch input.
Statistic
Value
Tasks / trajectories
2,512
Database instances
79
Average trajectory tokens
11,793.89
Median agent turns
16
Average tool calls
17.00
Average SQL executions
13.91
Appendix
Table 11: Statistics of the SFT corpus after teacher rollout and oracle-based selection.
Audit Check
Full Synthetic Pool vs. Spider2.0
Final SFT Subset vs. Spider2.0
Database Names
30 / 30 Identical
30 / 30 Identical
Database File SHA-256
30 / 30 Byte-identical
30 / 30 Byte-identical
Task ID Intersection
0
0
Exact Question Matches
0
0
Normalized Exact Matches
0
0
Same-DB Question Containment
0
0
Appendix
Table 12: Comprehensive leakage and overlap audit between synthetic training data and Spider 2.0-Lite SQLite-135 evaluation tasks.
Test ID
Nearest Synthetic ID
Database
Seq. Ratio
Task Divergence (Manual Verification)
local007
local_syn_000323
Baseball
0.1689
Test: Lifetime career span across all players. Train: Career home-run metrics for the 1990s debut cohort.
local310
local_syn_000250
F1
0.1688
Test: Annual minimal combinations of driver and constructor. Train: Constructor points share over 2010–2020.
local259
local_syn_000331
IPL
0.1425
Test: Lifetime batting and bowling profiles per player. Train: Venue-level toss-to-field match aggregations.
local055
local_syn_000045
Chinook
0.1419
Test: Customer spending by top vs. bottom sales artists. Train: Invoice revenue aggregations by billing country.
Table 13: Manual inspection of top-5 nearest train-test pairs by lexical sequence ratio within shared databases.
Metric (over 135 test tasks)
Median
P90
P95
Maximum
≥ 0.75
≥ 0.80
Document Cosine Similarity
0.4965
0.6229
0.6487
0.7122
0
0
Max-Chunk Cosine Similarity
0.5241
0.6306
0.6639
0.7075
0
0
Appendix
Table 14: Semantic similarity distributions between evaluation queries and candidate training queries within the same database.
Model
Protocol
S2-SQLite
BI-mini
LSB-SQLite
Qwen2.5-Coder-32B
Native
16.3%
3.0%
7.4%
DSH-SQL
3.0%
0%
5.6%
OmniSQL-32B
Native
14.8%
2.7%
8.9%
DSH-SQL
1.5%
0%
0%
Arctic-Text2SQL-7B
Native
15.6%
2.3%
7.0%
DSH-SQL
0%
0%
1.1%
Appendix
Table 15: Cross-harness evaluation of representative baselines. Native denotes the original or author-recommended inference protocol; DSH-SQL evaluates the same model in our interactive harness without additional harness-specific training.
Source
Verified trajectories
Share
Spider 2.0-Lite
1,797
71.5%
Spider 2.0-DBT
715
28.5%
Combined
2,512
100.0%
Appendix
Table 16: Verified trajectories retained after three teacher rollouts per task.