RareDx: Controlled Knowledge Integration and Graph-Grounded Policy Optimization for Rare-Disease Diagnosis
Organizations: Xi’an Jiaotong University · UIUC · Tsinghua University · University of Cambridge · Nankai University · Zhejiang University · Yale University
Abstract
Rare-disease diagnosis is a long-tail reasoning problem: phenotypes are incomplete, individual disorders are sparsely documented, and relevant evidence is distributed across ontologies, gene annotations, and biomedical text. Language models consequently favor common conditions, miss rare candidates, or produce plausible but invalid names. We introduce RareDx, which couples controlled evidence use with knowledge-graph-grounded policy optimization. RareDx-Harness normalizes heterogeneous records into one ranked-diagnosis task and compares direct inference, static retrieval, adaptive tools, and structured phenotype-gene-disease reasoning over a shared knowledge layer. The training pipeline combines Top-10 post-training with RareDx-KGPO, our knowledge-graph-grounded policy optimization method. Its reward projects predictions into a canonical disease graph and integrates curated graded relevance, ontology proximity, biomedical similarity, and phenotype consistency. Vocabulary and output-budget constraints prevent dense partial credit from rewarding fabricated or overlong differentials. Across eight benchmarks, the complete RareDx system centered on Qwen3.5-9B reaches 38.34 macro Hit@10, 1.60 points above GPT-5.5 under the archived protocol; a disjoint validation-selection audit retains a 6.80-point routing gain over Direct on held-out cases. The 27B system reaches 23.53/36.56/40.76 at Hit@1/5/10. Controlled ablations show that retrieval is not uniformly helpful and that controlled routing is central to the gain. These results indicate that structured medical knowledge can turn a compact model into a competitive diagnostic ranker across heterogeneous long-tail settings in clinical practice.
Figures & tables
| (a) Hit@1: first-ranked diagnosis | |||||||||
| Model | MyGene2 | RAMEDIS | MME | HMS | LIRICAL | RDS | RDC | Pheno. | Macro |
| Closed-source LLMs | |||||||||
| Claude Opus 4.7 | 9.6 | 21.1 | 12.5 | 22.2 | 23.4 | 12.4 | 15.2 | 25.6 | 17.8 |
| GLM-5.2 | 11.6 | 18.5 | 28.3 | 21.3 | 26.1 | 11.6 | 15.5 | 22.8 | 19.5 |
| GPT-5.5 | 20.5 | 21.5 | 29.5 | 23.8 | 27.2 | 10.7 | 12.1 | 30.8 | 22.0 |
| Open-weight baselines | |||||||||
| Panel | Configuration | Hit@1 | Hit@5 | Hit@10 |
|---|---|---|---|---|
| A | Direct | 17.03 | 26.63 | 31.27 |
| A | Adaptive ReAct | 16.49 | 27.51 | 33.53 |
| B | Direct | 17.2 | 23.7 | 27.5 |
| B | 3-hop, dense gene retrieval | 16.2 | 23.6 | 28.4 |
| B | 3-hop, HPO-Resnik | 22.2 | 31.7 | 36.3 |
| System | Hit@1 | Hit@5 | Hit@10 |
|---|---|---|---|
| Qwen3.8-27B, direct | 2 | 12 | 28 |
| Ours 9B, direct | 4 | 12 | 16 |
| Ours 9B, full Harness | 14 | 38 | 46 |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Cases | Dataset | Cases |
| MyGene2 | 146 | RAMEDIS | 624 |
| MME | 40 | HMS | 88 |
| LIRICAL | 370 | RDS | 1,803 |
| RDC | 678 | Phenopackets | 500 |
| Total | 4,249 | ||
| Model | Dataset | D, dev | A, dev | Pick | D, test | A, test | ||
|---|---|---|---|---|---|---|---|---|
| 9B | MyGene2 | 33 | 113 | 24.2/27.3/39.4 | 12.1/18.2/27.3 | D | 21.2/29.2/41.6 | 14.2/29.2/32.7 |
| 9B | RAMEDIS | 113 | 511 | 11.5/18.6/21.2 | 16.8/40.7/47.8 | A | 9.6/14.9/18.4 | 17.6/37.0/41.1 |
| 9B | MME | 9 | 31 | 0.0/11.1/11.1 | 0.0/0.0/0.0 | D | 6.5/9.7/9.7 | 12.9/12.9/12.9 |
| 9B | HMS | 22 | 66 | 22.7/27.3/27.3 | 27.3/45.5/59.1 | A | 9.1/12.1/13.6 | 24.2/37.9/47.0 |
| 9B | LIRICAL | 62 | 308 | 17.7/27.4/29.0 | 16.1/27.4/30.7 | A | 22.1/30.8/32.5 | 19.2/27.0/30.8 |
| 9B | RDS | 342 | 1,461 | 11.7/27.2/38.3 | – | D | 10.1/24.2/37.2 | – |
| Stage | Component | Rows | Content |
|---|---|---|---|
| SFT | train | 51,368 | RareArena cases with constructed Top-10 targets |
| SFT | validation | 932 | held-out Phenopacket-format prompts |
| RL | RareArena | 42,977 | article-backed case reports and test results |
| RL | Phenopacket | 8,392 | phenotype/gene disease prompts |
| RL | gene-phenotype hard set | 3,928 | synthetic hard cases |
| RL | simulated HPO subset | 2,714 | synthetic incomplete-phenotype cases |
| Reported row | Request alias | Interface | Decoding and tools |
|---|---|---|---|
| GPT-5.5 | gpt-5.5 | OpenAI-compatible proxy, xiaoai.plus/v1 | temperature, top- , seed, and output cap: provider default; tools: none |
| Claude Opus 4.7 | claude-opus-4-7 | OpenAI-compatible proxy, xiaoai.plus/v1 | temperature, top- , seed, and output cap: provider default; tools: none |
| GLM-5.2 | glm-5.2 | Alibaba MaaS compatible API for phenotype sets; DashScope Generation for RDS/RDC | decoding and output cap: provider default; tools: none |
| Method | Update | Signal | Teacher |
|---|---|---|---|
| GRPO | Group relative | Task reward | No |
| DAPO | Dynamic groups, asymmetric clip | Task reward | No |
| OPD | On-policy distillation | Token distributions | Yes |
| RareDx-KGPO (Ours) | Group relative | Medical graph reward | No |
| Model | Observed output |
|---|---|
| Qwen3.8-27B | Restates phenotype categories and surveys metabolic mechanisms, but does not produce a disease-ranked answer before truncation. MCADD is not recovered. |
| Ours 9B + Harness | Ranks MCADD first, followed by carnitine palmitoyltransferase II deficiency, very-long-chain acyl-CoA dehydrogenase deficiency, multiple acyl-CoA dehydrogenase deficiency, and carnitine-acylcarnitine translocase deficiency. |
| Direct model | Complete block | Mean words | Median words |
|---|---|---|---|
| Qwen3.8-27B | 0/50 | 462.2 | 451.0 |
| Ours 9B | 50/50 | 91.3 | 90.5 |
| Probe | Full | Ablated | Positive, full | Positive, ablated |
|---|---|---|---|---|
| Exact diagnosis | 1.000 | 1.000 | 100.0 | 100.0 |
| Phenotype neighbor | 0.316 | 0.008 | 100.0 | 0.8 |
| Pseudo-disease | -0.175 | 0.320 | 0.0 | 82.5 |
| Gold at rank 11 | 0.068 | 0.279 | 100.0 | 100.0 |
| Malformed output | 0.000 | 0.000 | 0.0 | 0.0 |
| Calls | Cases | Hit@1 | Hit@5 | Hit@10 |
|---|---|---|---|---|
| 0 | 324 | 2.16 | 3.40 | 4.01 |
| 1 | 383 | 35.77 | 54.31 | 57.96 |
| 2 | 1,423 | 38.65 | 54.25 | 59.31 |
| 3 | 4,100 | 22.10 | 37.90 | 43.32 |
| Strategy | LLM generations | Local retrieval | External API |
|---|---|---|---|
| Direct | 1 | None | No |
| Static RAG | 1 | Once | No |
| Adaptive ReAct | Adaptive | No | |
| Structured 3-hop | 1 | HPO + gene lookup | No |
| RRF fusion | Strategy-dependent | No |
| 27B | 35B | |||
|---|---|---|---|---|
| Extra lists | Hit@10 (%) | Output tokens | Hit@10 (%) | Output tokens |
| 0 | 34.38 | 173 | 32.81 | 179 |
| 1 | 36.72 | 346 | 32.81 | 354 |
| 2 | 39.06 | 519 | 36.72 | 529 |
| 4 | 39.84 | 864 | 42.19 | 877 |
| 8 | 46.09 | 1,556 | 42.19 | 1,575 |