cs.CLSep 28, 2026

Recursive LLM Degradation in Biomedical Question Answering: A Cross-Generation Study

Authors: Bibek Bhandari, Kshitij Lingthep

Organizations: Independent Researcher Memphis, Tennessee, USA

Abstract

Repeatedly training language models on their own generated data may create a synthetic-data feedback loop in which errors and distributional biases are reintroduced into subsequent training datasets. This paper studies that process in biomedical question answering (QA) using PubMedQA and two Qwen2.5 model sizes, 0.5B and 3B parameters. The study compares a recursive synthetic-data condition, in which generation G(k+1) is trained on answers produced by G(k), against a Human-Control condition that repeatedly uses the original human training data. The study evaluates across four generations from G0-G3 with two random seeds (42 and 123) and a fixed evaluation set of 1,000 expert-labeled samples. The evaluation includes disease and chemical entity F1, context-supported rate, lexical and semantic similarity, answer length, repetition rate, and other evaluation metrics. The Recursive condition for both model sizes and both seeds showed larger declines than the Human-Control condition in disease entity F1, chemical entity F1, context-supported rate, ROUGE-L, and cosine similarity. Under the fixed no-repeat 3-gram decoding constraint, the main observed behavioral change was increased answer length, while the measured 3-gram repetition rate did not increase. The magnitude of the difference-in-change was larger for the 3B model than for the 0.5B model. This difference was particularly apparent in disease F1, context-supported rate, cosine similarity, and answer length. These results show domain-specific changes associated with using recursive synthetic-data training in biomedical QA, but do not establish clinical hallucination rates or universal model collapse.

Figures & tables

Explore similar work

CardsList
  1. MedHopQA: A Disease-Centered Multi-Hop Reasoning Benchmark and Evaluation Framework for LLM-Based Biomedical Question Answering

    May 12, 2026Rezarta Islamaj, Robert Leaman, Joey Chan +13Biomedical TextClinical Reasoning Training

  2. Self-Study Reconsidered: The Hidden Fragility of Learning from Self-Generated QA

    Jun 30, 2026Ekaterina Alimaskina, Denis Shveykin, Gleb Molodtsov +3Faithful Question GenerationQuestion

  3. Domain Fine-Tuning vs. Retrieval-Augmented Generation for Medical Multiple-Choice Question Answering: A Controlled Comparison at the 4B-Parameter Scale

    Apr 26, 2026Avi-ad Avraam BuskilaMultiple-Choice QuestionsMedical Visual Question Answering