cs.CLSep 23, 2026

EviStreams: Human-in-the-Loop AI Data Extraction for Systematic Reviews in Medicine

Authors: Sai Karthik KosuriAnkita Shashikant BhosaleMichael GlickAlonso Carrasco-LabraChris Callison-Burch

Abstract

Systematic reviews underpin clinical guidelines, yet their data-extraction step is a major expert-labor bottleneck bound by a protocolized workflow: two reviewers extract each study independently, an adjudicator resolves disagreements, and the team keeps an auditable record of how every value was produced. Large language models can assist with extraction, but that assistance must fit established review protocols and preserve reproducibility. We present EviStreams, a live, open-source, no-code web platform that puts review teams in control of AI-assisted extraction at three key stages: program design (a structured decomposition approved before any code runs), field specification (typed field definitions calibrated from a pilot), and extracted predictions (reviewer-blinded dual review with adjudication). Working through a form builder, a domain expert defines typed fields rather than prompts, runs extraction over uploaded PDFs, inspects every value alongside the supporting passage it came from, and resolves a reviewer-blinded dual review into an auditable consensus export. An evaluation across four clinical corpora and three frontier model families, released with the system, shows that extraction quality is shaped far more by the field specification than by the choice of model. EviStreams is live at https://evistreams.com/demo and released under Apache-2.0.

Explore similar work

Apr 8, 2026cs.DL

TiAb Review Plugin: A Browser-Based Tool for AI-Assisted Study Selection in Systematic Reviews

Server-based screening tools impose subscription costs, while open-source alternatives require coding skills, and full-text screening has remained outside the scope of no-code open-source tools. We developed TiAb Review Plugin, an open-source Chrome browser extension that provides no-code, serverless artificial intelligence (AI)-assisted study selection covering both title and abstract (T&A) screening and full-text screening. It uses Google Sheets as a shared database and Google Drive as a PDF store, and users supply their own large language model (LLM) API key. For T&A screening, it offers manual review, LLM batch screening, and machine learning (ML) active learning. For full-text screening, it retrieves open-access PDFs from PubMed Central, Europe PMC, Unpaywall, OpenAlex, and publisher pages, supports blinded dual review with structured exclusion reasons and adjudication, optionally obtains an LLM judgment with page-anchored evidence, and computes PRISMA 2020 flow counts. We re-implemented the default ASReview algorithm (TF-IDF with Naive Bayes) in TypeScript and compared it with the Python original using 10-fold cross-validation on six datasets. For LLM T&A screening, we compared 16 parameter configurations on a benchmark dataset, validated the best (Gemini 3.0 Flash, low thinking budget, TopP 0.95) on five public datasets (1,038 to 5,628 records; 0.5% to 2.0% prevalence), and benchmarked nine further models from four developers. The TypeScript classifier produced top-100 rankings identical to ASReview on all six datasets. LLM T&A screening achieved recall of 94% to 100% with precision of 2% to 15%, and work saved over sampling at 95% recall (WSS@95) of 46.3% to 89.3%. No additional model exceeded the 96.1% recall of the reference configuration; the most recent models traded recall for precision. The classification accuracy of the full-text stage has not yet been evaluated.
Yuki Kataoka, Masahiro Banno, Michihito Kyo +8
Sep 14, 2026cs.SE

PRISMA-LLM: An Empirical Reporting Framework for AI-Assisted Systematic Reviews

Large language models (LLMs) and AI-enabled software increasingly participate in systematic-review decisions, yet the information needed to audit these workflows is reported inconsistently. We analyze SciLitBench, a corpus of 888 review-automation papers with 14,726 annotations, to characterize changes in methods, review-stage use, evaluation and reported limitations. Automation has shifted toward LLM- and software-facing workflows, including stages that can alter the evidence base. Since 2023, 38.0% of software/product papers reported no evaluation, compared with 9.3% of LLM papers. Reporting coverage increased with LLM workflow complexity, yet 52% of positive-only LLM evaluations still reported an unmet reliability or performance requirement. From these patterns, we introduce PRISMA-LLM, an empirically grounded framework separating implementation disclosure from consequence-sensitive evaluation and limitation reporting.
Miguel Zabaleta, Baihan Lin
Apr 24, 2026cs.CL

SLIDERS: Systematic Reviews via Automated Evidence Synthesis and Reconciliation

Systematic reviews -- which requires comprehensive evidence collection and synthesis from large document corpora in response to targeted research questions -- are foundational in finance, social sciences, and other technical fields. Manual construction of evidence tables is labor-intensive, and recent LLM-based assistants relying on embedding or keyword based search often fail to meet the coverage standards of systematic reviews. We introduce SLIDERS, a novel LLM-based methodology for systematic reviews, by automatically assembling evidence tables tailored to research questions. In addition to extracting structured data from documents, SLIDERS can extract full-text excerpts that serve as direct evidence or as provenance for structured data. Core to SLIDERS is an automated evidence reconciliation agent that writes code to analyze and reconcile extracted evidence, bringing together information fragmented across documents, resolving inconsistencies across excerpts, and synthesizing overlapping findings into a coherent evidence table. In addition, SLIDERS allows users to ask follow-up questions in natural language to further explore the assembled evidence. We evaluate SLIDERS on three systematic-review-style tasks over large document collections. SLIDERS outperforms the best-performing baseline across benchmarks, remains near 90% accuracy across 6M-11M-token corpora. On two new follow-up analysis benchmarks SLIDERS can answer 77.9% and 58.3% followup questions accurately
Harshit Joshi, Priyank Shethia, Jadelynn Dao +1