cs.CLOct 1, 2026

Evaluating Biomedical Reranking for LLM-Based Question Answering over Longitudinal Clinical Notes

Authors: Maryam Shahbaz Ali, Laura B. Strachan, Caitlin Sherman, Mark Kovler, Eleanor Mackey, Syed Muhammad Anwar

Organizations: Children’s National Hospital, Washington, DC, USA · University of Florida, Gainesville, FL, USA · George Washington University, Washington DC, USA

Abstract

Patient-specific clinical question answering requires locating the right evidence within long, heterogeneous longitudinal clinical records in which relevant facts may be scattered across encounters, repeated in copied-forward notes, or expressed using different clinical terminology. We evaluated whether biomedical reranking can improve evidence selection and downstream answer quality in a locally deployed retrieval-augmented generation pipeline for longitudinal clinical notes. The pipeline combines PubMedBERT dense retrieval, BM25 lexical retrieval, weighted reciprocal-rank fusion, and MedCPT cross-encoder reranking. Across 1,000 open- and closed-ended question-answer pairs from a cohort of 200 bariatric surgery patients, reranking increased exact source-chunk retrieval within the top 10 items, Hit@10 from 46.6% to 60.6% and mean reciprocal rank from 0.2371 to 0.3252. With Qwen3-8B generation, local judge-assessed answer correctness increased from 44.8% to 48.6%. These results show that biomedical reranking can improve the placement of relevant clinical evidence within a limited context window, although gains in retrieval do not translate proportionally into gains in answer correctness.

Explore similar work

CardsList
  1. When Retrieval Doesn't Help: A Large-Scale Study of Biomedical RAG

    Jun 2, 2026Erfan Nourbakhsh, Rocky Slavin, Ke Yang +1

  2. Retrieval Augmented Biomedical Question Answering with Weak Question Recovery and Neural Reranking for BioASQ Task 14b

    Aug 2, 2026Xueying Zhao, Lee Mai, Balaji AnandganeshCross-Encoder RerankingText Corpora