EviSearch: Trustworthy Extraction and Synthesis of Clinical Trial Evidence with Agents that Improve with Use
Organizations: Arizona State University · Mayo Clinic
Abstract
Structured extraction of evidence from clinical trial publications underpins systematic reviews and clinical guidelines, yet large language models are adopted for it only hesitantly: their outputs are difficult to verify, their use commonly requires transmitting documents to proprietary services, and they do not improve from the corrections their users make. We present EviSearch, a multi-agent system that addresses these three obstacles. Three tool-augmented agents with complementary access to a publication extract every column of an evidence table, and a value is admitted only after an attribution verifier has read it on its cited page, so that every value carries a page-level attribution. Disagreement between independent agents directs human review to the cells most likely to be wrong, and reviewer feedback refines the schema definitions and a curation knowledge base without updating model parameters. The agentic system runs entirely offline on open-weight models. On a clinician-annotated benchmark of randomized-trial publications, EviSearch attributes 100.0% of its values, reaches 91.70% accuracy autonomously, and reaches 95.22% after review of 15.6% of cells, exceeding random review of the strongest single agent at equal effort by 1.75 points.
Figures & tables
| Accuracy (%) | Attrib. (%) | ||||
| System | Run 1 | Run 2 | Mean | Cited | Verified |
| Single pass | 88.93 | 89.10 | 89.01 | 0.0 | 0.0 |
| PDF Query Agent (A) | 89.42 | 90.00 | 89.71 | 99.7 | 0.0 |
| Search Agent (B) | 92.16 | 92.37 | 92.27 | 100.0 | 0.0 |
| Oracle selection | 93.20 | 93.23 | 93.21 | 99.9 | 0.0 |
| EviSearch | 91.88 | 91.52 | 91.70 | 100.0 | 98.9 |
| Schema | Knowledge | Agent A | Agent B | EviSearch |
|---|---|---|---|---|
| Auto-drafted | guidelines | 88.35 | 90.86 | 89.91 |
| Reviewed 2 | guidelines | 89.31 | 91.80 | 90.49 |
| Reviewed 3 | guidelines | 89.44 | 91.23 | 91.04 |
| Reviewed 3 | + knowledge | 89.71 | 92.27 | 91.70 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Agreement- | Provenance- | |
| gated | gated | |
| Values dropped after a negative reading | 2.1 | 1.1 |
| of which correct (% of dropped) | 53 | 50 |
| Cells corrected by the reading | 0.26 | 0.38 |
| Cells made incorrect by the reading | 0.34 | 0.34 |
| Corroboration mark informative | no | yes |