Longitudinal clinical reasoning requires large language models (LLMs) to identify and integrate relevant evidence distributed across extended patient histories. Although long-context models can process increasingly large amounts of information, providing more history does not necessarily make relevant evidence more accessible or improve reasoning. We compare five context strategies (Full, Recent, Episodic, Semantic, and Hybrid) on MedLoCoMo across four open-weight LLMs, examining answer correctness, robustness to query-evidence distance, and abstention on questions with unsupported premises. Episodic and Hybrid generally achieve the strongest overall accuracy, while Recent Context degrades most as supporting evidence becomes more distant; Episodic and Hybrid maintain the highest accuracy at long distances. Analysis of adversarial questions further shows that strong performance on answerable questions does not necessarily translate to successful abstention when the available history does not support the requested conclusion. These findings show that reliable longitudinal reasoning depends not only on how much history an LLM can access, but critically on how relevant evidence is selected and presented for reasoning.
Figures & tables
Fig. 1: Overview of the experimental framework. Given a longitudinal patient history and query, a context strategy determines the information provided to the answer model. We compare Full Context, Recent Context, Episodic, Semantic, and Hybrid strategies under the same underlying histories and queries.
Question Type
Scope
Count
%
Medical reasoning
Single
142
14.2
Care-plan rationale
Single
143
14.3
Adversarial
Single
143
14.3
Longitudinal progression
Cross
143
14.3
Cross-admission comparison
Cross
143
14.3
Frequency pattern
Cross
143
14.3
TABLE I: Distribution of question types in the 1,000-question evaluation subset.
Statistic
Sessions
Turns
Est. Tokens
Mean
29.5
1,662.3
35,072.6
Median
28.0
1,623.0
34,099.0
Min
17
913
20,451
Max
64
3,522
74,255
TABLE II: Longitudinal context statistics for the patient timelines. Token counts are estimated using the same approximation employed for experimental evidence budgeting.
Fig. 2: Context strategies evaluated for longitudinal clinical reasoning. Full and Recent provide direct-context baselines, while Episodic, Semantic, and Hybrid use relevance-based retrieval over encounter-level evidence, patient-specific facts, or their combination, respectively.
Fig. 3: Illustrative example of the five context strategies.
Fig. 4: Overall accuracy and judge sensitivity on MedLoCoMo. (a–b) End-to-end accuracy across context strategies and answer models using Gemma4 and MedGemma as correctness judges, respectively; error bars indicate 95% Wilson confidence intervals. (c) Agreement between the two correctness judges, including overall agreement, Cohen’s κ , and directional disagreements. (d) BERTScore F1 between generated and reference answers for the primary Gemma4-evaluated runs; percentages within cells indicate the proportion of generations evaluable by BERTScore.
Strategy
Coverage (%)
Precision (%)
Recall (%)
F1 (%)
Full Context
44.9
22.0
27.6
21.2
Recent Context
63.2
13.5
11.1
11.2
Episodic
87.5
48.7
50.0
46.1
Semantic
79.0
35.7
40.5
34.8
Hybrid
87.9
41.3
47.9
40.4
TABLE III: Session-level citation grounding across context strategies. Coverage denotes the percentage of evaluated responses containing at least one citation that could be mapped to a source session.
Model
Strategy
CA-Adv
CA-Comp
CA-Freq
CA-Long
SA-Adv
SA-Care
SA-Med
Qwen-3-8B
Full Context
6.99
51.05
41.96
53.85
26.57
66.43
71.13
Recent Context
0.70
45.45
38.46
31.47
16.78
46.15
42.25
Episodic
1.40
55.24
49.65
51.75
30.07
88.11
76.76
Semantic
0.00
48.95
34.27
51.05
37.06
72.73
61.97
Hybrid (E+S)
2.80
56.64
47.55
60.14
36.36
86.71
78.17
Mistral-7B
Full Context
32.17
11.19
1.40
1.40
2.10
16.08
4.93
TABLE IV: Accuracy (%) by context strategy and question type on the MedLoCoMo dataset across four answer models. Each model is evaluated on 1,000 questions: Cross-Admission Adversarial (CA-Adv; n=143 ), Cross-Admission Comparison (CA-Comp; n=143 ), Cross-Admission Frequency Pattern (CA-Freq; n=143 ), Cross-Admission Longitudinal Progression (CA-Long; n=143 ), Single-Admission Adversarial (SA-Adv; n=143 ), Single-Admission Care-Plan Rationale (SA-Care; n=143 ), and Single-Admission Medical Reasoning (SA-Med; n=142 ). Best performance within each answer model and question type is shown in bold.
Fig. 5: Accuracy across evidence-distance bins for each context strategy and answer model. Evidence distance is measured as the session gap between the query and its supporting evidence. Points indicate accuracy within each distance bin, with error bars showing 95% Wilson confidence intervals.
Strategy
Model
Answerable
Adversarial
Full
Qwen
56.9
16.8
Mistral
7.0
17.1
LLaMA
52.1
12.9
GLM-4
63.3
12.2
Mean
44.8
14.8
Recent
Qwen
40.8
8.7
TABLE V: Accuracy (%) on answerable and adversarial questions. Adversarial questions contain unsupported premises or request conclusions that cannot be established from the available patient history. Mean denotes the average across the four answer models. Bold indicates the highest model-specific accuracy within each context strategy and question category.
Fig. 6: End-to-end judged accuracy on LongMemEval across context strategies and answer models. Error bars indicate 95% Wilson confidence intervals.