cs.AIOct 6, 2026

MedZERO: Self-Evolving Agents for Open-Ended Medical Reasoning Through Controlled Knowledge Accumulation

Authors: Xilin Dang, Weilin Ruan, Xue Yang, Jinghao Wang, Xiaowei Hu, Jinpeng Li, Pheng-Ann Heng

Abstract

Large language models (LLMs) have shown promise in medical question answering and clinical reasoning, yet their improvement remains constrained by static parametric knowledge and costly expert supervision. Self-evolving agents offer a promising alternative by enabling models to improve through iterative task generation and problem-solving. However, most existing self-evolving methods are designed for easily verifiable domains such as mathematics and coding, where solutions can be checked by exact answers or executable programs. Medical reasoning is fundamentally different: it is open-ended, knowledge-intensive, and often only partially verifiable. We present MedZERO, a self-evolving framework for open-ended medical reasoning. MedZERO couples an Examiner that generates frontier medical question-option pairs with a Reasoner that solves them through evidence-grounded multi-turn reasoning with external knowledge tools. To support reliable, continual improvement, MedZERO adopts controlled knowledge accumulation, which maintains temporary exploratory knowledge and curated persistent knowledge in reasoning. We evaluate MedZERO on five public medical reasoning benchmarks using 4B- and 8B-scale base models under open-ended evaluation. Across all settings, MedZERO consistently outperforms the underlying base models and prior self-evolving baselines, achieving up to 13.7 average accuracy-point gains over the next-best self-evolving baseline.

Explore similar work

CardsList
  1. SEMA-RAG: A Self-Evolving Multi-Agent Retrieval-Augmented Generation Framework for Medical Reasoning

    May 16, 2026Yongfeng Huang, Ruiying Chen, James ChengRetrieval-Augmented GenerationHealthcare

  2. EMR: Self-Evolving Medical Multi-Agent System via Experience Mining and Reuse

    Sep 14, 2026Dongsheng Shi, Yue Li, Xin Yi +1Continual Learning for LLM AgentsMulti-Agent LLM Systems

  3. Evo-PI: Aligning Medical Reasoning via Evolving Principle-Guided Supervision

    Jun 30, 2026Xianda Zheng, Huan Gao, Meng-Fen Chiang +3LLM AlignmentMedical VLMs