cs.CLSep 16, 2026

A Benchmark Framework for Screening Automation in Systematic Reviews

Authors: Gauransh Kumar, Luciano Marchezan, Guillaume Genois, Kévin Delcourt, Eugene Syriani

Organizations: Université de Montréal Montreal, Canada

Abstract

Systematic reviews (SR) are essential for evidence-based research, but their screening phase is highly time-consuming and labor-intensive. Large language models (LLMs) offer a promising opportunity to reduce this workload by assisting with article relevance classification. However, existing evaluation approaches often rely on traditional metrics that may be misleading for highly imbalanced SR screening datasets. This paper presents a benchmark dataset of 45 06445\,064 labeled entries for evaluating LLM performance in SR screening across 32 curated secondary studies. It proposes an evaluation framework that accounts for class imbalance, i.e., the natural prevalence of excluded articles relative to included articles in SRs. It also introduces PromptSR, a tool designed to support prompt experimentation, experiment management, and result analysis for LLM-based screening. We also present a use case demonstrating the application of SRBench and PromptSR.

Figures & tables

Explore similar work

CardsList
  1. Beyond Accuracy: LLM Variability in Evidence Screening for Software Engineering SLRs

    Apr 29, 2026Gilberto Sussumu Hida, Danilo Monteiro Ribeiro, Erika YahataSystematic ReviewScreening

  2. Understanding LLMs in Title-Abstract Screening: From Disagreements to Recommendations

    Jun 16, 2026Mika Mäntylä, Patricia Matsubara, Katia Romero Felizardo +5Systematic ReviewPeer Review

  3. PRISMA-LLM: An Empirical Reporting Framework for AI-Assisted Systematic Reviews

    Sep 14, 2026Miguel Zabaleta, Baihan LinSystematic ReviewPeer Review