From Given to Gathered Evidence: Agentic Learning for Longitudinal Medical Reasoning
Organizations: Durham University · DAMO Academy, Alibaba Group · Hupan Laboratory · University of the Chinese Academy of Sciences
Abstract
Foundation models can serve as clinical agents through tool-use harnesses. However, conventional medical benchmarks assess reasoning over preselected evidence rather than the ability to seek it across clinical records and longitudinal imaging. We propose CASE: a series of role-specific Clinical Agents for Seeking Evidence, together with a tool-use harness and an agentic post-training framework for compact vision-language policy models. We further introduce a longitudinal multimodal benchmark built on UK Biobank, comprising 50,401 clinical questions derived from real-world ICD-10-coded diagnoses of 4,739 participants. Each question links to a patient-specific environment containing clinical context and multi-sequence MRI from baseline and follow-up visits, where agents autonomously select which visits, organs, modalities, slices, and specialist tools to inspect and compare. Supervised fine-tuning transfers evidence-seeking workflows from 14,734 frontier-model interaction trajectories, followed by agentic reinforcement learning on the learner's own environment interactions. Privileged on-policy self-distillation and rubric-based LLM feedback refine evidence-to-conclusion reasoning without prescribing tool sequences. Experiments show that CASE moves beyond question-answer imitation toward transferable investigation policies, strengthening evidence-grounded longitudinal reasoning. Under matched evaluation conditions, our Qwen3-VL-8B based agent achieves over 16% and 10% relative improvements in answer accuracy over GPT-5.4 and Claude Opus 4.8. Code will be available at https://github.com/VinyehShaw/CASE.
Figures & tables
| Domain | Oncology | Endocrine/ metabolic | Cardiovascular | Respiratory/ allergy | Multisystem | Medical history | Overall (%) | |||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Method / Question | Cancer (B) | Diabetes (B) | Condition (S) | Condition set (M) | Hyper- tension (B) | Condition (S) | Condition set (M) | ICD-10 (S) | Disease systems (M) | History (B) | Micro | Macro |
| Base model | ||||||||||||
| Qwen3-VL-8B | 67.7 | 46.1 | 72.9 | 70.0 | 38.4 | 52.5 | 49.3 | 26.9 | 5.8 | 41.1 | 46.0 | 47.1 |
| Medical models | ||||||||||||
| Lingshu-32B | 85.0 | 70.4 | 55.0 | 54.6 | 48.8 | 58.9 | 51.6 | 35.2 | 2.9 | 40.5 | 49.9 | 50.3 |
| Lingshu-7B | 65.4 | 83.1 | 48.6 | 40.3 | 65.8 | 50.7 | 38.6 | 28.5 | 3.3 | 42.8 | 45.5 | 46.7 |
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
| Clinical area | Target | Form | Questions |
|---|---|---|---|
| Oncology | Doctor-diagnosed cancer | B | 4,710 |
| Endocrine/metabolic | Doctor-diagnosed diabetes | B | 4,718 |
| Cardiovascular | Condition identification | S | 4,685 |
| Cardiovascular | Condition set | M | 4,739 |
| Cardiovascular | Essential hypertension (I10) | B | 4,037 |
| Respiratory/allergy | Condition identification | S | 5,013 |
| Paradigm | Benchmark input | Model behavior | Evaluation target |
|---|---|---|---|
| Medical QA | Question and task-provided text or images; external knowledge access where permitted. | Infer an answer from available context, optionally retrieving external knowledge. | Primarily answer correctness. |
| Interactive medical tasks | Task, patient state or records, and tools or simulated encounters. | Acquire information and perform task-dependent decisions or actions. | Task-specific success; some benchmarks also assess reasoning or information acquisition. |
| CASE | Clinical question and a participant-specific longitudinal evidence environment. | Select tools, accumulate text and images, synthesize findings, and answer. | Exact-set accuracy; logged retrieval and completion metrics. Independent grounding assessment is an additional validation protocol. |
| System | Evidence use and interaction | Reported adaptation mechanism |
|---|---|---|
| MDAgents ( Kim et al., 2024 ) | Medical problem solving with complexity-dependent individual or collaborative consultation. | Inference-time selection of roles and collaboration structure. |
| MedRAX ( Fallahpour et al., 2025 ) | Iterative chest X-ray analysis using specialist tools. | Tool orchestration without additional agent training. |
| MedAgentSim ( Almansoori et al., 2025 ) | Doctor–patient dialogue and requested tests or imaging results. | Retrieval of successful encounters and reflections from corrected failures, combined with multi-agent reasoning. |
| MedChain-Agent ( Liu et al., 2025 ) | Information gathering and sequential referral, examination, diagnosis, and treatment. | Iterative agent feedback and MedCase-RAG with an expanding case database. |
| ClinicalAgent ( Yan et al., 2025 ) | Department routing and staged consultation with laboratory and imaging reports. | Performance-informed clinician assignment and collaborative synthesis. |
| MedAgent-Pro ( Wang et al., 2026b ) | Guideline-derived plans, patient-specific tool execution, and quantitative clinical indicators. | Hierarchical reasoning with stepwise evidence checks. |
| System | Parameter-learning mechanism | Relation to the present study |
|---|---|---|
| MMedAgent ( Li et al., 2024 ) | Instruction tuning of tool calls and answers across medical tasks, including composed tool use. | Establishes medical tool-policy learning through supervised examples. |
| MedVR ( Jiang et al., 2026b ) | RL with uncertainty-guided visual exploration and rollout-consensus supervision. | Learns image inspection; correctness gates its auxiliary tool reward. |
| Ophiuchus ( Jiang et al., 2026a ) | Tool-oriented SFT, reflection fine-tuning, and RL over interleaved visual interactions. | Establishes learning beyond demonstrations for medical visual tools. |
| MACRO ( Fan et al., 2026 ) | Supervised initialization and GRPO encouraging discovered composite-tool use. | Combines parameter learning with experience memory and tool-set expansion. |
| CASE | Demonstration initialization followed by learner-directed RL with an insight-focused, detached self-distillation penalty. | Studies acquisition across longitudinal patient sources and synthesis feedback combined with outcome and rubric rewards before group normalization. |
| Configuration | Tools | Returned evidence | Call budget |
|---|---|---|---|
| Anatomist | vista3d_segment ; vista3d_compare_stations ; vista3d_compare_visits | Organ/station measurements; longitudinal changes; segmentation overlays. | 4 |
| Radiologist | list_available_images ; view_image ; compare_visits | Acquisition availability and metadata; selected or paired MRI views. | 6 |
| Consultant | All six above; list_available_reports ; get_report ; recall_skills ; write_skill | Combined patient evidence, generated report claims, and reusable textual guidance. | 12 |
| Step | Content |
|---|---|
| Question | Based on longitudinal medical imaging and clinical records, which hepatic conditions are supported by the available evidence? Select all that apply. Options: (A) hepatic steatosis (fatty liver disease); (B) benign hepatic lesion; (C) primary liver malignancy; (D) liver cirrhosis; (E) none of the above. |
| System | You are a clinical consultant agent, the final arbitration over heterogeneous clinical evidence rather than any single source. You command the full ten-tool Consultant registry of table 7 . To inspect imaging you may list available acquisitions ( list_available_images ), view a selected slice or montage ( view_image ), and compare matched views across visits ( compare_visits ); for quantitative structure you may segment an organ ( vista3d_segment ) and compare volumes across stations ( vista3d_compare_stations ) or across visits ( vista3d_compare_visits ); for textual claims you may list the generated MedGemma reports ( list_available_reports ) and read one via get_report(organ, report_type) with report_type in findings, impression, or followup for liver, pancreas, or wholebody, treating these draft reports as claims to verify because they may hallucinate; and for accumulated experience you may recall reusable skills ( recall_skills ), recording new ones ( write_skill ) only in the later review stage. Required workflow: (1) call recall_skills for the implicated organ(s) first; (2) review evidence from multiple sources (images, volumes, reports), using each for what it is best at; (3) actively doubt—mark key claims [ verified ], [ suspicious ], or [ contradicted ] and verify every decision-relevant suspicious claim with a direct tool call; (4) reflect—build the causal chain and explicitly consider what could mislead you here. When done calling tools, the final message (no tool calls) must contain all five headers in order, under 2000 characters: Phase 1 Skill Recall & Review Plan; Phase 2 Multi-source Review; Phase 3 Doubt & Verification; Phase 4 Reflective Insight (the causal chain from verified evidence toward the answer, plus what could have misled and why it was rejected); Phase 5 Final Answer. Rules: maximum 12 tool calls (hard-enforced), and once about 9 are used stop calling tools and report with the evidence at hand; prefer high-value calls ( recall_skills first, vista3d_compare_visits for change questions, get_report for report evidence, view_image only for targeted verification) and skip sources irrelevant to this question; never copy a report conclusion without verification when a quantitative tool can check it; the Answer: line in Phase 5 is mandatory and the answer is extracted only from it; do not call write_skill at this stage (skill recording happens in a later review stage). |
| User (patient info) | Patient ID: 2306330. This female participant was born in 1955 and was 64 years of age at the time of the most recent assessment (Instance 3). Longitudinal follow-up data were available from two assessment visits: baseline assessment (Instance 2) was conducted on 2017-08-16 at age 61 years, and follow-up assessment (Instance 3) on 2019-12-16 at age 64 years, a follow-up interval of 2.3 years. Regarding housing tenure, the participant transitioned from own with a mortgage (Instance 0) to own outright (Instance 3), indicating a change in housing stability, and resided in a house or bungalow. Household composition at Instance 3 comprised 2 resident(s); the household reported ownership of three vehicle(s), a change from two at Instance 2. The reported average total household income before tax at Instance 3 was less than 18,000, compared to 52,000–100,000 at Instance 2, suggesting a potential decrease in household economic resources. Regarding occupational characteristics, the participant reported commuting distance of nan units from home to workplace (changed from 13 units at Instance 2) and employment duration of nan units in their primary occupation. This comprehensive sociodemographic profile provides essential context for interpreting clinical and imaging findings, as socioeconomic factors are well-established determinants of health outcomes and disease risk in population-based cohort studies. |
| P1 Plan | The case contains liver abnormalities including a reported focal lesion. I begin by recalling liver-specific skills, then select the sources best suited to each sub-question. The main clinical challenge is distinguishing an incidental or benign hepatic lesion from primary liver malignancy. I will not directly convert an imaging abnormality into a cancer diagnosis. Instead, I will characterize lesion morphology, evaluate longitudinal stability, quantify liver structural changes, and verify whether clinical evidence supports malignancy. |
| Tool call 1 | recall_skills(organ=liver; keywords=hepatic lesion characterization, hepatocellular carcinoma, hepatic steatosis, malignancy verification) Liver lesion assessment requires correlation between morphology, interval change, and clinical confirmation. Stable lesions without invasive characteristics are less suggestive of malignancy. Imaging suspicion alone should not be interpreted as confirmed cancer. call_6b57153c6e84494a94af5fa8 |
| Tool call 2 | list_available_images() Available imaging (Visit 1/Visit 2 pairs): liver—IDEAL multiecho and IDEAL temporal; whole body—Dixon BH 17s W (stations 4/8/12/20/24); pancreas—gradient-echo multiecho. call_26553c9a231a453ea5b4c90 |
| Setting | RSFT | Agentic RL defaults |
|---|---|---|
| Training data and epochs | Mixed three-role; one epoch | Mixed three-role; one epoch |
| Learning rate | ||
| Batch | devices samples accumulation | prompts rollouts |
| Sequence budget | 8,192 tokens | 4,096 prompt; 6,144 continuation; 12,288 total |
| Assistant-turn budget | Recorded conversation | 10 |
| Image handling | Resize parameter 512 | Backend-specific cached images |
| Variant | Phase-4 synthesis for participant 1588571 |
|---|---|
| GRPO + RLAIF (no OPSD) | The pancreatic imaging findings offer only indirect evidence relevant to diabetes, including fat deposition, atrophy, or secondary complications. However, these findings are not diagnostic of diabetes, these findings are not diagnostic of diabetes, these findings are not diagnostic of diabetes… |
| CASE-8B (full) | The current MRI findings provide no direct basis for diagnosing diabetes, although pancreatic fat, atrophy, or associated complications may offer supporting context. In this case, the reports describe preserved anatomy and improving pancreatic fat content, without diabetes-related complications or a documented diagnosis in the clinical history. Therefore, the available evidence does not support selecting the diabetes option. |
| Method | Teacher / privileged information | How feedback enters learning |
|---|---|---|
| GKD ( Agarwal et al., 2024 ) | External teacher on student-generated sequences | Distribution matching on learner-visited prefixes; compatible with joint RL fine-tuning. |
| RLAD ( Zhang et al., 2026c ) | External teacher and old-student mixture | Mixture-anchored policy ratios make imitation selective according to rollout advantage. |
| TGPO ( Liu et al., 2026 ) | External teacher at student prefixes | Cross entropy on teacher-preferred next tokens supplements trajectory-level verifiable-reward optimization. |
| OPDVR ( Lin et al., 2026 ) | External teacher and outcome verifier | Outcome correctness gates the direction of sampled-token distillation rewards within a policy-gradient objective. |
| TRACE ( Wang et al., 2026a ) | Synchronized self-teacher with coarse diagnostic labels | Differentiable student distillation is routed to annotated reasoning spans; other positions retain GRPO. The default emphasizes forward KL on key spans of correct rollouts. |
| SSOPD ( Tan and Hong, 2026 ) | Self-generated correct completion | A forward-KL auxiliary loss supervises early prefixes of an unsuccessful completion; GRPO remains the outcome objective. |
| Criterion | Allocation | Assessment |
|---|---|---|
| Evidence identification | Specific findings rather than generic abnormality claims. | |
| Logical reasoning | A coherent connection between findings, interpretation, and answer. | |
| Answer support | Evidence that supports the selected conclusion. | |
| Completeness | The intermediate steps needed to justify the conclusion. |
| Domain | Oncology | Endocrine/ metabolic | Cardiovascular | Respiratory/ allergy | Multisystem | Medical history | Overall (%) | |||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Method / Question | Cancer (B) | Diabetes (B) | Condition (S) | Condition set (M) | Hyper- tension (B) | Condition (S) | Condition set (M) | ICD-10 (S) | Disease systems (M) | History (B) | Micro | Macro |
| (a) Anatomist Agent | ||||||||||||
| Qwen3-VL-8B-Instruct | 67.6 | 27.6 | 78.7 | 69.9 | 36.5 | 43.4 | 38.5 | 8.8 | 5.9 | 45.3 | 39.5 | 42.2 |
| Lingshu-32B | 89.4 | 82.9 | 98.5 | 97.8 | 72.4 | 65.5 | 68.7 | 17.7 | 0.6 | 28.4 | 59.2 | 62.2 |
| Lingshu-7B | 88.0 | 95.2 | 54.3 | 35.2 | 78.0 | 54.5 | 35.4 | 13.6 | 2.8 | 44.6 | 47.1 | 50.2 |
| MedGemma-27B | 41.2 | 25.5 | 19.2 | 46.9 | 23.7 | 13.0 | 14.6 | 7.9 | 2.5 | 27.9 | 21.0 | 22.2 |
| Method | Micro (%) 95% CI | Full method (pp) 95% CI | Relative gain (%) 95% CI |
|---|---|---|---|
| (a) Anatomist Agent | |||
| CASE (full) | 77.5 4.3 | — | — |
| CASE (RSFT only) | 64.9 3.9 | 12.6 1.8 | 19.4 3.8 |
| Qwen3-VL-8B-Instruct | 39.5 3.4 | 38.0 2.3 | 96.2 13.5 |
| Qwen3.8-Max | 65.8 4.0 | 11.7 2.1 | 17.8 3.8 |
| GPT-5.4-0305 | 63.6 4.0 | 13.9 2.0 | 21.9 3.7 |
| Training variant | Micro (%) | Insight quality | Judge coverage (%) |
|---|---|---|---|
| (a) Anatomist Agent Shared judge coverage: 84.1% | |||
| RSFT only | 64.9 | 0.45 | 94.2 |
| GRPO (correctness + format) | 73.0 | 0.48 | 90.6 |
| GRPO + RLAIF | 74.3 | 0.65 | 85.3 |
| GRPO + OPSD | 77.1 | 0.59 | 93.1 |
| CASE-8B | 77.5 | 0.70 | 95.0 |
| Evidence condition | RSFT micro | Full micro | Full macro | Full drop (pp) | Full calls |
|---|---|---|---|---|---|
| Question only | 35.8 | 38.6 | 36.9 | 37.5 | 0 |
| Initial clinical context | 48.4 | 52.7 | 50.9 | 23.4 | 0 |
| Fixed evidence packet | 62.2 | 66.0 | 64.4 | 10.1 | 0 |
| Adaptive acquisition, all sources | 65.7 | 76.1 | 75.1 | — | 5.3 |
| Adaptive, without MRI views | 60.9 | 65.9 | 65.1 | 10.2 | 4.7 |
| Adaptive, earlier visit only | 59.8 | 64.6 | 62.9 | 11.5 | 3.9 |
| Policy | Calls | Image blocks | Visual positions | Assistant tokens | Latency (s) median / p95 | Failure (%) |
|---|---|---|---|---|---|---|
| CASE-8B (full) | 5.1 | 8.4 | 3612 | 1452 | 34.2 / 68.5 | 2.6 |
| CASE-8B (RSFT only) | 4.6 | 7.2 | 3087 | 1248 | 31.5 / 61.2 | 6.8 |
| Qwen3-VL-8B-Instruct | 2.9 | 3.6 | 1514 | 782 | 22.4 / 49.1 | 21.4 |
| Qwen3.8-Max | 5.3 | 8.6 | 3704 | 1603 | 43.6 / 86.2 | 4.9 |
| GPT-5.4-0305 | 4.3 | 7.0 | 3012 | 1324 | 46.1 / 92.4 | 5.7 |
| Claude Opus 4.8 | 4.1 | 6.8 | 2896 | 1281 | 44.3 / 89.0 | 5.2 |
| Forward | Context | Parameters | Role in learning |
|---|---|---|---|
| Ordinary scoring | Student history | Current snapshot, fixed | Produces ; no scoring gradient. |
| Privileged scoring | Same history plus | Same snapshot, fixed | Produces ; no scoring gradient. |
| Actor | Ordinary history | Trainable policy | Optimizes sampled assistant actions. |
| Reference | Ordinary history | Frozen SFT policy | Supplies the separate reference penalty. |