Large Language Models for Structured Clinical Data Analysis: Dual-Agent Grounding and Validation
Authors: Erfan D. Dehkalani, Seetha Shankaran, Abbot R. Laptook, C. Michael Cotten, P. Ellen Grant, Yangming Ou
Organizations: Fetal-Neonatal Neuroimaging and Developmental Science Center, Boston Children’s Hospital, and Harvard Medical School, Boston, MA, USA · Department of Pediatrics, Wayne State University School of Medicine, Detroit, MI, USA · Department of Pediatrics, Women & Infants Hospital of Rhode Island, and Warren Alpert Medical School of Brown University, Providence, RI, USA · Division of Neonatology, Department of Pediatrics, Duke University School of Medicine, Durham, NC, USA
Objective: To develop and characterize CLEAR-Med, a dual-agent framework for natural-language analysis of structured clinical data that separates SQL-based invocation from independent validation. Methods: CLEAR-Med uses one agent to translate a question into executable Structured Query Language (SQL), retain the executed query and database result, and produce a draft. Deterministic checks and a separately invoked cross-provider Validation Agent then accept the draft, request one bounded repair, or abstain. We formalized the system as a bounded selective pipeline and evaluated CLEAR-Med's configuration and scalability, and the Invocation Agent's accuracy and consistency on a 25-query development benchmark, using a harmonized 21-site neonatal hypoxic-ischemic encephalopathy table containing 532 de-identified infant records and approximately 1,300 variables. Results: CLEAR-Med completed all six nominal scalability configurations, including 500x1300. Across 25 development-benchmark queries repeated five times, the Invocation Agent answered 83 of 125 responses correctly (66.4%; query-cluster bootstrap 95% CI, 48.0-83.2%), compared with 15 of 125 (12.0%; 95% CI, 3.2-22.4%) for the ungrounded ChatGPT baseline, a paired improvement of 54.4 percentage points (95% CI, 36.8-72.0%). Conclusion: CLEAR-Med provides a general architecture for traceable analysis of structured clinical data: numerical claims remain linked to executed SQL, and unresolved cases can fail closed. The reported experiments characterize CLEAR-Med's configuration and scalability and the Invocation Agent's accuracy, while the formal analysis establishes the encoded-property guarantee of the complete control flow; a prospective full-pipeline evaluation of the validation and abstention stages is the next stage of this work.
Figures & tables
Problem or Issue
Natural-language analysis of governed clinical tables can produce fluent numerical answers without exposing the cohort, calculation, or supporting rows.
What is Already Known
Language models can generate SQL and use external tools, but query generation alone does not establish that a reported number preserves the executed result and intended clinical constraints.
What this Paper Adds
CLEAR-Med separates schema-guided SQL invocation from deterministic checks and an independently invoked Validation Agent. The pipeline retains SQL provenance, permits one bounded repair, abstains on unresolved failures, and provides a property-specific encoded-validity guarantee.
Who would benefit from the new knowledge in this paper
Biomedical-informatics researchers and clinical-data teams building auditable natural-language interfaces to governed relational data.
Table 1 : Statement of Significance
System
50 × 50
50 × 100
100 × 50
100 × 100
50 × 1000
500 × 1300
Max Time (s)
Notes
OpenBioLLM-8B [ 6 ]
✓
✓
✓
✓
×
×
2.19
CUDA OOM
open-bio-med-8B [ 8 ]
✓
✓
✓
✓
×
×
11.78
CUDA OOM
Med-ChimeraLlama-3-8B [ 7 ]
✓
✓
✓
✓
×
×
26.05
CUDA OOM
Daredevil-8B [ 22 ]
✓
✓
✓
✓
×
×
26.00
CUDA OOM
DISC-MedLLM [ 11 ]
✓
×
×
×
×
×
32.86
Timeout
CLEAR-Med
✓
✓
✓
✓
✓
✓
28.12
–
Table 2 : End-to-end scalability across six fixed table configurations under the deployment interfaces described in Section 3.9. The table reports deterministic completion status and elapsed time for each tested configuration.
Condition
Err (%) ↓
5%-GT (%) ↑
SD (%) ↓
Range (%) ↓
ChatGPT baseline (OpenAI API)
55.74
6.67
58.80
146.58
Invocation Agent
33.67
65.00
35.07
78.55
Table 3 : Accuracy and consistency on numerical queries across repeated trials ( K=5 repetitions, 12 percentage-valued queries). Err = mean absolute percentage error vs. the deterministic data oracle; 5%-GT = share of responses within 5% of ground truth; SD and Range are computed across repetitions and averaged over queries. Arrows ( ↓ / ↑ ) indicate the preferred direction.
Task type
n /condition
ChatGPT baseline
Invocation Agent
comparison
30
0.0
30.0
count
25
36.0
100.0
multi-condition
20
5.0
50.0
odds ratio
10
20.0
100.0
percentage
40
7.5
72.5
Overall (95% CI)
125
12.0 [3.2, 22.4]
66.4 [48.0, 83.2]
Table 4 : Per-task accuracy across 25 queries × 5 repetitions. n is the number of responses per condition in each subgroup. Overall accuracy ( n=125 per condition) is reported with a query-cluster bootstrap 95% confidence interval.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Metric
Operational definition
Goal
Time-based diagnostics
Setup environment time
Time to initialize the computational environment
↓
Load files to database
Time to load data files into the database
↓
Setup database time
Time to configure the database for query processing
↓
Setup language model time
Time to initialize the language model
↓
Create agent time
Time to instantiate the SQL agent
↓
Appendix
Table A.1 : Supporting definitions and results. Panel A: operational definitions for the efficiency and execution diagnostics used to compare ChatGPT configurations. The final column gives the preferred direction; an em dash denotes a descriptive metric. All time measures are in seconds.
Metric
Operational definition
Goal
Performance diagnostics
Execution time
Duration from invocation start to final response
↓
Total time
Execution time plus all setup operations
↓
Number of database invocations
Count of successful database connections and queries
—
Number of errors
Count of SQL syntax errors and execution failures
↓
Query diagnostics
Appendix
Table A.1 : Supporting definitions and results (continued). Panel C: diagnostics used in the one-factor-at-a-time configuration screen. The final column gives the preferred direction where applicable; an em dash denotes a descriptive diagnostic.
Factor
Setting
Exec. (s)
Total (s)
Errors
DB inv.
SQL len.
SQL compl.
Selected
Max iterations
10
317.12
317.30
–
3
503
5
50
134.19
134.71
–
1
555
6
✓
100
139.08
139.28
–
1
208
2
Max execution time
10 s
38.52
38.71
3
0
–
–
30 s
49.90
50.13
3
0
–
–
60 s
85.40
85.60
3
1
–
–
✓
Appendix
Table A.1 : Supporting definitions and results (continued). Panel D: one-factor-at-a-time configuration screen for CLEAR-Med. The check mark identifies the setting selected for the reported configuration. Selection first required successful SQL generation and execution, then used errors and execution cost; SQL length and complexity are descriptive. An em dash indicates that the diagnostic was not recorded for that factor.
Difficulty
n /condition
ChatGPT baseline
Invocation Agent
easy
45
24.4
86.7
medium
25
8.0
80.0
hard
55
3.6
43.6
Appendix
Table A.1 : Supporting definitions and results (continued). Panel E: accuracy by query difficulty. n is the number of responses per condition; the clustered interval for the overall result appears in Table 4.
While interpretable prototype networks offer compelling case-based reasoning for clinical diagnostics, their raw continuous outputs lack the semantic structure required for medical documentation. Bridging this gap via standard Retrieval-Augmented Generation (RAG) routinely triggers ``retrieval sycophancy,'' where Large Language Models (LLMs) hallucinate post-hoc rationalizations to align with visual predictions. We introduce ProtoMedAgent, a framework that formalizes multimodal clinical reporting as an iterative, zero-gradient test-time optimization problem over a strict neuro-symbolic bottleneck. Operating on a frozen prototype backbone, we distill latent visual and tabular features into a discrete semantic memory. Online generation is strictly constrained by exact set-theoretic differentials and a reflective Scribe-Critic loop, mathematically precluding unsupported narrative claims. To safely bound data disclosure, we introduce a semantic privacy gate governed by k-anonymity and ℓ-diversity. Evaluated on a 4,160-patient clinical cohort, ProtoMedAgent achieves 91.2% Comparison Set Faithfulness where it fundamentally outperforms standard RAG (46.2%). ProtoMedAgent additionally leverages a binding ℓ-diversity phase transition to systematically reduce artifact-level membership inference risks by an absolute 9.8%.
Clinical data-science agents must transform heterogeneous longitudinal records into auditable analyses, yet existing benchmarks largely isolate medical question answering, structured-table reasoning, or generic scientific repositories. We introduce CLINLENS, a benchmark of 200 executable tasks over five linked MIMIC resources spanning structured electronic health records, notes, electrocardiograms, chest radiographs, and echocardiograms. A 4 x 5 taxonomy crosses four patient-time scopes with five analysis capabilities. Program-first reverse synthesis pairs each bounded semi-raw package with an evaluator-private reference workflow and checks required artifacts, cohort and temporal semantics, and the final answer. On a fixed 126-task suite, the strongest of 24 standardized model-scaffold configurations achieves 56.3% scope-macro STRICTPASS despite 100% EXECSUCCESS. For reference, a separately configured coding agent solves 83 of 126 tasks, while five biomedical systems adapted to GPT-4o-mini reach at most 2.9% scope-macro STRICTPASS. These results expose a substantial gap between runnable submissions and correct clinical analyses.
Deploying Large Language Models (LLMs) in high-stakes clinical settings remains limited by structural hallucinations, weak deterministic reasoning over tabular patient data, and omissions in vector retrieval. This paper presents the architecture and validation of Medi-Gemma, a Clinical Decision Support System (CDSS) for wound pathology triage and workflow automation. The platform introduces a decoupled framework that separates clinical perception from data orchestration while preserving traceable reasoning. Medi-Gemma uses a multi-stage pipeline coordinated by a centralized ClinicalOrchestrator. Data requests are handled without generative inference by a DataManager that cleans unstructured Electronic Medical Record (EMR) files through type coercion. Natural language queries are processed by a hierarchical IntentRouter, which routes requests to deterministic analytics paths executed by a PandasQueryEngine or to patient-specific reasoning managed by a ClinicalRAGEngine using a CPU-optimized vector store. A key contribution is the Ground Truth Injection Module, which intercepts patient-specific queries, extracts numeric identification tokens, queries the structured dataframe via Pandas, retrieves the latest validated clinical state, and embeds this snapshot as an overriding context block in the LLM prompt before generation. Safety compliance is enforced by a deterministic ProtocolManager that maps clinical terminology to fixed evidence-based risk pathways, while a SafetyVerifier phrase filter prevents output rule violations. Validation shows that this architecture eliminates semantic context drift, prevents database compilation crashes, and improves factual adherence to backend clinical repositories. These results support Medi-Gemma as a safer pattern for LLM-based clinical decision support where structured data fidelity, retrieval grounding, and deterministic safeguards are essential.
Mohammed Saim Ahmed Quadri, Yunzhe Xue, Justin W. Ady +1
Department of Computer Science, New Jersey Institute of Technology Newark, NJ, USA · Department of Data Science, New Jersey Institute of Technology Newark, NJ, USA · Vascular and Endovascular Surgery, Robert Wood Johnson Hospital New Brunswick, NJ, USA