Large Language Models for Structured Clinical Data Analysis: Dual-Agent Grounding and Validation
Authors: Erfan D. Dehkalani, Seetha Shankaran, Abbot R. Laptook, C. Michael Cotten, P. Ellen Grant, Yangming Ou
Organizations: Fetal-Neonatal Neuroimaging and Developmental Science Center, Boston Children’s Hospital, and Harvard Medical School, Boston, MA, USA · Department of Pediatrics, Wayne State University School of Medicine, Detroit, MI, USA · Department of Pediatrics, Women & Infants Hospital of Rhode Island, and Warren Alpert Medical School of Brown University, Providence, RI, USA · Division of Neonatology, Department of Pediatrics, Duke University School of Medicine, Durham, NC, USA
Objective: To develop and characterize CLEAR-Med, a dual-agent framework for natural-language analysis of structured clinical data that separates SQL-based invocation from independent validation. Methods: CLEAR-Med uses one agent to translate a question into executable Structured Query Language (SQL), retain the executed query and database result, and produce a draft. Deterministic checks and a separately invoked cross-provider Validation Agent then accept the draft, request one bounded repair, or abstain. We formalized the system as a bounded selective pipeline and evaluated CLEAR-Med's configuration and scalability, and the Invocation Agent's accuracy and consistency on a 25-query development benchmark, using a harmonized 21-site neonatal hypoxic-ischemic encephalopathy table containing 532 de-identified infant records and approximately 1,300 variables. Results: CLEAR-Med completed all six nominal scalability configurations, including 500x1300. Across 25 development-benchmark queries repeated five times, the Invocation Agent answered 83 of 125 responses correctly (66.4%; query-cluster bootstrap 95% CI, 48.0-83.2%), compared with 15 of 125 (12.0%; 95% CI, 3.2-22.4%) for the ungrounded ChatGPT baseline, a paired improvement of 54.4 percentage points (95% CI, 36.8-72.0%). Conclusion: CLEAR-Med provides a general architecture for traceable analysis of structured clinical data: numerical claims remain linked to executed SQL, and unresolved cases can fail closed. The reported experiments characterize CLEAR-Med's configuration and scalability and the Invocation Agent's accuracy, while the formal analysis establishes the encoded-property guarantee of the complete control flow; a prospective full-pipeline evaluation of the validation and abstention stages is the next stage of this work.
Figures & tables
Problem or Issue
Natural-language analysis of governed clinical tables can produce fluent numerical answers without exposing the cohort, calculation, or supporting rows.
What is Already Known
Language models can generate SQL and use external tools, but query generation alone does not establish that a reported number preserves the executed result and intended clinical constraints.
What this Paper Adds
CLEAR-Med separates schema-guided SQL invocation from deterministic checks and an independently invoked Validation Agent. The pipeline retains SQL provenance, permits one bounded repair, abstains on unresolved failures, and provides a property-specific encoded-validity guarantee.
Who would benefit from the new knowledge in this paper
Biomedical-informatics researchers and clinical-data teams building auditable natural-language interfaces to governed relational data.
Table 1 : Statement of Significance
System
50 × 50
50 × 100
100 × 50
100 × 100
50 × 1000
500 × 1300
Max Time (s)
Notes
OpenBioLLM-8B [ 6 ]
✓
✓
✓
✓
×
×
2.19
CUDA OOM
open-bio-med-8B [ 8 ]
✓
✓
✓
✓
×
×
11.78
CUDA OOM
Med-ChimeraLlama-3-8B [ 7 ]
✓
✓
✓
✓
×
×
26.05
CUDA OOM
Daredevil-8B [ 22 ]
✓
✓
✓
✓
×
×
26.00
CUDA OOM
DISC-MedLLM [ 11 ]
✓
×
×
×
×
×
32.86
Timeout
CLEAR-Med
✓
✓
✓
✓
✓
✓
28.12
–
Table 2 : End-to-end scalability across six fixed table configurations under the deployment interfaces described in Section 3.9. The table reports deterministic completion status and elapsed time for each tested configuration.
Condition
Err (%) ↓
5%-GT (%) ↑
SD (%) ↓
Range (%) ↓
ChatGPT baseline (OpenAI API)
55.74
6.67
58.80
146.58
Invocation Agent
33.67
65.00
35.07
78.55
Table 3 : Accuracy and consistency on numerical queries across repeated trials ( K=5 repetitions, 12 percentage-valued queries). Err = mean absolute percentage error vs. the deterministic data oracle; 5%-GT = share of responses within 5% of ground truth; SD and Range are computed across repetitions and averaged over queries. Arrows ( ↓ / ↑ ) indicate the preferred direction.
Task type
n /condition
ChatGPT baseline
Invocation Agent
comparison
30
0.0
30.0
count
25
36.0
100.0
multi-condition
20
5.0
50.0
odds ratio
10
20.0
100.0
percentage
40
7.5
72.5
Overall (95% CI)
125
12.0 [3.2, 22.4]
66.4 [48.0, 83.2]
Table 4 : Per-task accuracy across 25 queries × 5 repetitions. n is the number of responses per condition in each subgroup. Overall accuracy ( n=125 per condition) is reported with a query-cluster bootstrap 95% confidence interval.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Metric
Operational definition
Goal
Time-based diagnostics
Setup environment time
Time to initialize the computational environment
↓
Load files to database
Time to load data files into the database
↓
Setup database time
Time to configure the database for query processing
↓
Setup language model time
Time to initialize the language model
↓
Create agent time
Time to instantiate the SQL agent
↓
Appendix
Table A.1 : Supporting definitions and results. Panel A: operational definitions for the efficiency and execution diagnostics used to compare ChatGPT configurations. The final column gives the preferred direction; an em dash denotes a descriptive metric. All time measures are in seconds.
Metric
Operational definition
Goal
Performance diagnostics
Execution time
Duration from invocation start to final response
↓
Total time
Execution time plus all setup operations
↓
Number of database invocations
Count of successful database connections and queries
—
Number of errors
Count of SQL syntax errors and execution failures
↓
Query diagnostics
Appendix
Table A.1 : Supporting definitions and results (continued). Panel C: diagnostics used in the one-factor-at-a-time configuration screen. The final column gives the preferred direction where applicable; an em dash denotes a descriptive diagnostic.
Factor
Setting
Exec. (s)
Total (s)
Errors
DB inv.
SQL len.
SQL compl.
Selected
Max iterations
10
317.12
317.30
–
3
503
5
50
134.19
134.71
–
1
555
6
✓
100
139.08
139.28
–
1
208
2
Max execution time
10 s
38.52
38.71
3
0
–
–
30 s
49.90
50.13
3
0
–
–
60 s
85.40
85.60
3
1
–
–
✓
Appendix
Table A.1 : Supporting definitions and results (continued). Panel D: one-factor-at-a-time configuration screen for CLEAR-Med. The check mark identifies the setting selected for the reported configuration. Selection first required successful SQL generation and execution, then used errors and execution cost; SQL length and complexity are descriptive. An em dash indicates that the diagnostic was not recorded for that factor.
Difficulty
n /condition
ChatGPT baseline
Invocation Agent
easy
45
24.4
86.7
medium
25
8.0
80.0
hard
55
3.6
43.6
Appendix
Table A.1 : Supporting definitions and results (continued). Panel E: accuracy by query difficulty. n is the number of responses per condition; the clustered interval for the overall result appears in Table 4.
Department of Computer Science, New Jersey Institute of Technology Newark, NJ, USA · Department of Data Science, New Jersey Institute of Technology Newark, NJ, USA · Vascular and Endovascular Surgery, Robert Wood Johnson Hospital New Brunswick, NJ, USA