Coverage-Aware Reasoning with Medical Tokens for Diagnosis Prediction
Organizations: Fuzhou University · Shanghai Jiao Tong University · Beihang University
Abstract
Large language models (LLMs) offer promising potential for next-visit diagnosis prediction, owing to their ability to integrate longitudinal clinical evidence and reason over it in natural language. However, reinforcement learning for LLM reasoning commonly rewards each trajectory according to the correctness of its final answer. In next-visit diagnosis prediction, multiple diagnoses can be simultaneously valid, but independently rewarding one diagnosis per trajectory does not distinguish repeated hits from coverage of different diagnoses. The policy can therefore concentrate on a few correct diagnoses, leaving others uncovered. Meanwhile, LLM tokenizers can split ICD codes into several generic tokens with limited clinical meaning, requiring multiple decoding steps to predict each diagnosis and hindering reasoning over a large disease vocabulary. To address both challenges, we propose CARing, a framework that represents diagnoses with compositional Semantic IDs (SIDs) and optimizes reasoning trajectories for multi-label coverage. Concretely, we first encode ontology-enriched disease semantics into compact SIDs through residual quantization, and ground the resulting SID tokens in natural language and longitudinal EHR contexts through multi-task alignment and reasoning-enriched training to unlock transferable LLM reasoning. CARing further improves unordered multi-label prediction through a coverage reward for reinforcement learning and multi-positive supervision. At inference time, the model supports both efficient direct constrained decoding and multi-chain reasoning with rank fusion. On MIMIC-III and MIMIC-IV, CARing exceeds all EHR-trained baselines in weighted F1 and attains the highest top-k recall at every reported cutoff, including R@30 of 46.04% and 46.52% in reasoning mode. Our codes and logs are available at https://github.com/zmlxzyh/CARing-Codes-Logs.
Figures & tables
| Model | MIMIC-III | MIMIC-IV | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| w- | R@10 | R@20 | R@30 | R@40 | w- | R@10 | R@20 | R@30 | R@40 | |
| RETAIN | 21.25 | 24.13 | 34.30 | 40.12 | 45.91 | 19.54 | 25.42 | 33.82 | 38.91 | 42.78 |
| KAME | 18.58 | 22.55 | 31.43 | 37.67 | 42.54 | 10.59 | 17.97 | 25.01 | 29.56 | 33.15 |
| StageNet | 18.13 | 21.36 | 30.37 | 36.64 | 41.10 | 11.46 | 18.41 | 25.29 | 30.20 | 33.82 |
| CGL | 21.79 | 25.51 | 35.28 | 41.52 | 46.75 | 20.08 | 24.07 | 33.96 | 40.04 | 44.31 |
| TRANS | 19.10 | 23.32 | 32.26 | 38.39 | 43.14 | 15.84 | 20.55 | 28.71 | 34.37 | 38.67 |
| Metric | Layer | Name only | Name + ICD path |
|---|---|---|---|
| R@30 (%) | 38.97 | 39.73 | |
| Codebook utilization (%) | 1 | 100.00 | 100.00 |
| 2 | 98.83 | 100.00 | |
| 3 | 97.27 | 100.00 | |
| Token distribution entropy | 1 | 7.943 | 7.951 |
| 2 | 7.542 | 7.919 |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | MIMIC-III | MIMIC-IV |
|---|---|---|
| # of patients | 5,449 | 79,393 |
| # of visits | 14,141 | 408,990 |
| Avg. # of visits per patient | 2.60 | 5.15 |
| Max. # of visits per patient | 29 | 170 |
| # of unique diagnoses | 3,874 | 37,917 |
| # of CCS codes | 285 | 842 |
| Metric | Name only | Name + ICD path |
| First-token groups pure by ICD / 256 | 58 | 143 |
| First-token groups pure by ICD / 256 | 108 | 212 |
| Top-level ICD weighted purity (%) | 69.36 | 90.51 |
| Same ICD group given shared token (%) | 57.60 | 86.49 |
| First-token groups pure by CCS / 256 | 13 | 44 |
| First-token groups pure by CCS / 256 | 36 | 103 |
| Model | IFEval | GSM8K |
|---|---|---|
| Base Qwen3-1.7B | 0.6820 | 0.8832 |
| S2 = S1 + Enriched Alignment | 0.1460 | 0.5238 |
| S3 = S2 + General Reasoning | 0.4713 | 0.6300 |