ARCagent: An Adaptive Retrieval Calibration Agent for Clinical Question Answering
Abstract
In diseases where clinical guidelines are incomplete, contested, or mutually contradictory, knowledge completeness and dynamic conflict-aware synthesis are two safety-critical properties that standard Retrieval-Augmented Generation systems do not provide. Therefore, we present \sysname, an adaptive retrieval calibration clinical question-answering agent for ME/CFS, a disease where diagnostic frameworks coexist and major guidelines actively contradict each other on treatment. ARCagent contributes three components. First, a 1,706-chunk, 10-source knowledge base with a structured inter-guideline conflict registry spanning all active ME/CFS diagnostic frameworks. Second, a conflict-aware retrieval calibration pipeline that re-ranks retrieved evidence using query-specific focus and conflict signals. Third, a benchmark scored by LLM-as-Judge, avoiding systematic underestimation averaging 10.1 percentage points caused by keyword matching. ARCagent achieves 95.3%, outperforming all base LLMs. Code is available at https://github.com/Yukyin/ARCagent.
Figures & tables
| Source | Organisation | Chunks | Words |
| IOM/NAM 2015 | US NAM | 929 | 109,145 |
| CDC ME/CFS | US CDC | 215 | 20,529 |
| NICE NG206 | UK NICE | 214 | 24,285 |
| Clinician Coalition | ME/CFS CC, USA | 102 | 35,700 |
| IACFS/ME Primer | Friedberg et al. | 82 | 24,464 |
| IQWiG N21-01 | IQWiG, Germany | 58 | 17,219 |
| Model | Auto | Diag | Navi | PEM | Scop | Conf | Spec | RevH | Over |
|---|---|---|---|---|---|---|---|---|---|
| Domain-specific medical LLMs | |||||||||
| BioMedLM | 76.3 | 77.8 | 76.7 | 72.6 | 42.6 | 60.8 | 73.5 | 68.3 | 68.6 |
| HuatuoGPT-II | 83.7 | 85.2 | 80.6 | 76.7 | 50.3 | 69.4 | 77.1 | 76.2 | 74.9 |
| Med-PaLM 2 | 85.1 | 83.5 | 82.7 | 81.2 | 54.6 | 73.1 | 80.4 | 81.7 | 77.8 |
| MedAlpaca | 70.8 | 73.2 | 71.3 | 69.4 | 39.5 | 54.6 | 71.2 | 63.3 | 64.2 |
| General-purpose LLMs | |||||||||
| Model | LLM | Keyword | Human |
|---|---|---|---|
| Domain-specific medical LLMs | |||
| BioMedLM | 68.6 | 55.2 | 65.2 |
| HuatuoGPT-II | 74.9 | 60.1 | 75.3 |
| Med-PaLM 2 | 77.8 | 61.3 | 75.4 |
| MedAlpaca | 64.2 | 52.8 | 65.1 |
| General-purpose LLMs | |||
| Configuration | Chun | Auto | Diag | Navi | PEM | Scop | Conf | Spec | RevH | Overall |
|---|---|---|---|---|---|---|---|---|---|---|
| Knowledge type ablation | ||||||||||
| w/o guideline knowledge | 18 | 95.8 | 95.4 | 96.1 | 95.3 | 77.3 | 88.5 | 94.2 | 92.8 | 91.9 |
| w/o empirical knowledge | 1,688 | 98.3 | 97.8 | 97.7 | 98.1 | 80.1 | 90.5 | 95.7 | 94.6 | 94.1 |
| Source-level ablation | ||||||||||
| w/o diagnostic frameworks | 689 | 96.1 | 95.8 | 96.6 | 95.7 | 79.8 | 89.2 | 94.5 | 93.4 | 92.6 |
| w/o treatment guidelines | 1,492 | 97.2 | 96.5 | 97.3 | 97.4 | 77.6 | 91.5 | 95.1 | 95.8 | 93.6 |
| System | Response excerpt | Conflict surfaced |
|---|---|---|
| Query. “Is graded exercise therapy still recommended for ME/CFS?” (clinician mode) | ||
| Grok-4.20 | No, graded exercise therapy (GET) is not recommended for ME/CFS by major clinical guidelines. UK NICE Guidelines explicitly advise against GET, defining it as fixed incremental activity increases from a baseline, and state that such programs should not be offered due to reports of harm and lack of evidence for benefit, recommending pacing within energy limits to avoid post-exertional malaise (PEM) instead. US CDC similarly notes that standard exercise recommendations for healthy people can be harmful for ME/CFS patients and recommends pacing to manage energy and prevent PEM flare-ups, reflecting a broader expert shift away from the old deconditioning model toward recognising ME/CFS as a biological illness with PEM as a core feature. | No |
| ARCagent | Graded exercise therapy is no longer recommended for ME/CFS under current major guidelines. NICE NG206 [1] explicitly contraindicates GET on the grounds that the deconditioning model underpinning it has not been validated and that progressive exercise can provoke post-exertional malaise. Conflict detected. IQWiG N21-01 [2] reports low-certainty short-term benefit for structured activity in some subgroups, diverging from NICE NG206 based on a different appraisal of the PACE trial. Current North American and UK consensus supports individualised activity pacing within each patient’s energy envelope. [1][3] | Yes, NICE NG206 vs IQWiG N21-01 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.