Large language models (LLMs) offer promising potential for next-visit diagnosis prediction, owing to their ability to integrate longitudinal clinical evidence and reason over it in natural language. However, reinforcement learning for LLM reasoning commonly rewards each trajectory according to the correctness of its final answer. In next-visit diagnosis prediction, multiple diagnoses can be simultaneously valid, but independently rewarding one diagnosis per trajectory does not distinguish repeated hits from coverage of different diagnoses. The policy can therefore concentrate on a few correct diagnoses, leaving others uncovered. Meanwhile, LLM tokenizers can split ICD codes into several generic tokens with limited clinical meaning, requiring multiple decoding steps to predict each diagnosis and hindering reasoning over a large disease vocabulary. To address both challenges, we propose CARing, a framework that represents diagnoses with compositional Semantic IDs (SIDs) and optimizes reasoning trajectories for multi-label coverage. Concretely, we first encode ontology-enriched disease semantics into compact SIDs through residual quantization, and ground the resulting SID tokens in natural language and longitudinal EHR contexts through multi-task alignment and reasoning-enriched training to unlock transferable LLM reasoning. CARing further improves unordered multi-label prediction through a coverage reward for reinforcement learning and multi-positive supervision. At inference time, the model supports both efficient direct constrained decoding and multi-chain reasoning with rank fusion. On MIMIC-III and MIMIC-IV, CARing exceeds all EHR-trained baselines in weighted F1 and attains the highest top-k recall at every reported cutoff, including R@30 of 46.04% and 46.52% in reasoning mode. Our codes and logs are available at https://github.com/zmlxzyh/CARing-Codes-Logs.
Figures & tables
Figure 1: Motivation for multi-label reinforcement learning in next-visit diagnosis prediction. Standard RL can collapse onto one correct diagnosis, leaving others uncovered. Our coverage reward discounts repeated predictions and encourages coverage of distinct correct diagnoses.
Figure 2: Overview of CARing. The framework constructs compositional medical SIDs, aligns them with clinical language and reasoning, combines a coverage reward with multi-positive answer supervision, and supports direct and multi-chain reasoning inference.
Model
MIMIC-III
MIMIC-IV
w- F1
R@10
R@20
R@30
R@40
w- F1
R@10
R@20
R@30
R@40
RETAIN
21.25
24.13
34.30
40.12
45.91
19.54
25.42
33.82
38.91
42.78
KAME
18.58
22.55
31.43
37.67
42.54
10.59
17.97
25.01
29.56
33.15
StageNet
18.13
21.36
30.37
36.64
41.10
11.46
18.41
25.29
30.20
33.82
CGL
21.79
25.51
35.28
41.52
46.75
20.08
24.07
33.96
40.04
44.31
TRANS
19.10
23.32
32.26
38.39
43.14
15.84
20.55
28.71
34.37
38.67
Table 1: Overall next-visit diagnosis prediction performance on MIMIC-III and MIMIC-IV. All values are percentages. Best and second-best results are shown in bold and underlined, respectively.
Metric
Layer
Name only
Name + ICD path
R@30 (%)
38.97
39.73
Codebook utilization (%)
1
100.00
100.00
2
98.83
100.00
3
97.27
100.00
Token distribution entropy
1
7.943
7.951
2
7.542
7.919
Table 2: SID quality and MIMIC-III recall with disease names alone or names plus ICD ontology paths. Entropy is in bits; its maximum is 8 for a 256-entry codebook.
Figure 3: Effect of the four stages of reasoning-enriched training on MIMIC-III under non-thinking direct inference.
Figure 4: Training ablation on MIMIC-III in thinking mode.
Figure 5: MIMIC-III trajectories for S5 (EM reward) and S6 (coverage reward), both initialized from S4.
Figure 6: MIMIC-III trajectories for S6 and S7. The missing S7 value at step 42 is linearly interpolated between steps 41 and 43.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
MIMIC-III
MIMIC-IV
# of patients
5,449
79,393
# of visits
14,141
408,990
Avg. # of visits per patient
2.60
5.15
Max. # of visits per patient
29
170
# of unique diagnoses
3,874
37,917
# of CCS codes
285
842
Appendix
Table 3: Dataset statistics after cohort preprocessing and before subsampling MIMIC-IV.
Figure 7: ICD ontology subgroup counts for four randomly sampled first SID tokens. Subgroups are the second nodes below the ontology root.
Metric
Name only
Name + ICD path
First-token groups pure by ICD / 256
58
143
First-token groups ≥80% pure by ICD / 256
108
212
Top-level ICD weighted purity (%)
69.36
90.51
Same ICD group given shared token (%)
57.60
86.49
First-token groups pure by CCS / 256
13
44
First-token groups ≥80% pure by CCS / 256
36
103
Appendix
Table 4: Semantic organization of first SID tokens over 4,491 distinct diseases. ICD groups are the first nodes below the ontology root.
Model
IFEval
GSM8K
Base Qwen3-1.7B
0.6820
0.8832
S2 = S1 + Enriched Alignment
0.1460
0.5238
S3 = S2 + General Reasoning
0.4713
0.6300
Appendix
Table 5: General-ability evaluation across training stages. Higher scores are better.