MedKIT: Evaluating Knowledge Integration and Generalization in Large Language Models
Authors: Lukas Thede, Yash Kumar Atri, David Chen, Danielle Bitterman, Matthias Bethge, Tom Hartvigsen, Zeynep Akata
Organizations: University of Tübingen, Tübingen AI Center · Helmholtz Munich · Munich Center for Machine Learning (MCML) · University of Virginia · University of British Columbia · Harvard Medical School · Technical University of Munich
Constantly evolving real-world knowledge necessitates models to be updated continuously. Especially in medicine, as clinical evidence changes over time, outdated knowledge can pose safety risks. Existing evaluations of knowledge integration focus on factual recall, offering limited insight into whether newly integrated knowledge is actually usable. Our benchmark MedKIT (Medical Knowledge Integration and Transfer) provides a granular evaluation of how models integrate and apply knowledge under realistic sequences of clinical updates. Each instance corresponds to a factual update derived from clinical evidence, paired with targeted probes that assess transfer across lexical variation, relational transformations, compositional reasoning, and open-ended operationalization, as well as locality tests for knowledge preservation. Using MedKIT, we conduct a large-scale empirical study of 12 knowledge integration strategies across 5 diverse models, including both general-purpose and medical LLMs. Our results reveal a consistent gap between recall and usable knowledge: while most methods achieve strong gains on the original update task and under lexical variation, relational generalization is limited, and no method yields meaningful improvements on compositional or operational tasks. These findings highlight a fundamental challenge in knowledge integration and position MedKIT as a testbed for developing methods that make newly integrated knowledge more consistently usable across tasks and contexts.
Figures & tables
Figure 1 : MedKIT construction pipeline and examples. Structured clinical comparisons are converted into factual updates and multiple task variants probing lexical, relational, compositional, and operational generalization, alongside locality tests for knowledge preservation.
Benchmark
Domain
Input
Generalization
L
R
C
O
CounterFact [ 30 ]
G
Triples
✓
-
-
-
ZSRE [ 26 ]
G
QA
✓
-
-
-
WikiBigEdit [ 39 ]
G
QA / Wiki facts
✓
✓
✓
-
MQuAKE [ 54 ]
G
QA (multi-hop)
-
✓
-
-
DocTER [ 47 ]
G
Documents
-
-
-
-
Table 1 : Comparing factual knowledge integration benchmarks. Prior work emphasizes factual recall and isolated generalization settings, with limited coverage of multi-level transfer, open-ended tasks, and realistic update scenarios. (G: General, M: Medical, L: Lexical, R: Relational, C: Compositional, O: Operational)
Figure 2 : MedKIT statistics and temporal structure. The benchmark spans 6,196 clinically grounded factual updates across conditions and oncology groups, with balanced labels and temporally ordered batches that enable realistic sequential evaluation.
Figure 3 : Generalization of integrated knowledge across task formats. Mean post–pre performance ( Δ ) across tasks from direct recall ( Update ) to increasingly demanding generalization settings. Top: aggregated by method; bottom: aggregated by model, with faint markers indicating individual model–method combinations. Across both views, performance consistently degrades beyond lexical variation, indicating that improvements in recall do not translate into reliable generalization across reformulated and open-ended task settings.
Figure 4 : Retention under sequential updates. Current Δ measures immediate gains, Previous Δ performance on earlier updates. Retention varies by method, reflecting differences in underlying integration mechanisms.
Figure 5 : Capability preservation under knowledge integration. Left: Update success versus locality change, measuring whether methods integrate target updates while preserving neighboring oncology facts. Right: Changes in general capabilities, grouped into latent competence, behavioral preferences, and protocol compliance, highlight the risk of updates interfering with the model.
Method
Update
Lexical
Compos.
Locality
GRACE
+75.5
+0.0
+0.0
+0.0
MEMIT
+38.1
+22.8
0.0
-1.9
MEMOIR
+52.2
+38.1
-0.8
-2.2
SEEKR
+41.9
+20.2
-1.9
-0.9
Table 2: Cross-domain evaluation on WikiBigEdit. Mean post–pre change in exact-match containment (pp) across two models.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Strategy
# batches
Median size
Mean size
Max size
Daily
98
2
2.9
17
Weekly
48
4
5.8
21
Monthly
14
16.5
20.2
44
Appendix
Table 3: Batch-size statistics for the post-2025 evaluation window.
Judge
OQ κ ( k=3 )
OQ acc
OG κ ( k=2 )
OG acc
OG top-1
claude-opus-4
1.00
100 %
1.00
100 %
✓
claude-3-5-haiku
1.00
100 %
1.00
100 %
✓
gemini-2.0-flash
1.00
100 %
1.00
100 %
✓
gemini-2.5-pro
1.00
100 %
1.00
100 %
✓
gpt-4o-mini
1.00
100 %
1.00
100 %
✓
gpt-4o
1.00
100 %
0.42
86.7 %
✓
Appendix
Table 4: Cross-judge tier-classification agreement. Per-judge: rank-based LOJO Cohen’s κ and accuracy against the 6-judge consensus tier labels. The final column reports whether the judge identifies Oracle-RAG-abs as the strongest-performing OG method. Bottom row: Fleiss’ κ across all 7 judges’ independently-derived tier labels. k=3 on OQ (oracle / middle-performing / degradation), k=2 on OG (oracle / remaining methods).
Figure 6 : Consensus per-method judge scores (mean over 7 judges) with 95% tag-resampling bootstrap confidence intervals and tier bands. Left: compositional task (OQ, k=3 ). Right: operational task (OG, k=2 ). On OQ, oracle retrieval methods form a distinct top-performing tier, while AlphaEdit forms a separate degradation regime. On OG, the data support only a stable oracle-vs-rest separation, with the remaining methods clustering near pre-update performance.
Judge
OQ ρ
OQ κ
OQ MAE
OG ρ
OG κ
OG MAE
OG pref.
claude-opus-4
0.86
0.81
0.45
0.54
0.48
0.66
57.5%
gpt-4o
0.83
0.78
0.52
0.64
0.57
0.59
51.0%
claude-3-5-haiku
0.84
0.80
0.43
0.42
0.29
0.84
47.9%
llama-3.3-70b
0.84
0.80
0.46
0.45
0.34
0.92
53.0%
gemini-2.0-flash
0.77
0.73
0.54
0.54
0.48
0.66
52.5%
gemini-2.5-pro
0.77
0.79
0.57
0.64
0.52
0.72
59.5%
Appendix
Table 5: Agreement with the two-clinician consensus on the 100-item annotation set. For each judge we report Spearman correlation ( ρ ), quadratic-weighted Cohen’s κ , and mean absolute error (MAE) on the 1–5 scale against the mean of the two independent clinician annotators, plus operational-task agreement on the categorical preference flag (OG pref.), averaged over the two clinicians. The final row reports inter-clinician agreement as a human agreement reference.
Figure 7 : Pairwise Spearman ρ between the seven candidate judges and two independent clinicians on the 100-item annotation set. Left: compositional OQ. Right: operational OG. Inter-clinician agreement is ρ=0.76 on OQ and ρ=0.62 on OG; judge agreement is generally higher on the structured compositional task than on open-ended operational generation.
Model
Domain
Scale
HuggingFace ID
Gemma-3-4B-IT
General
4 B
google/gemma-3-4b-it
Qwen-3-4B-Instruct
General
4 B
Qwen/Qwen3-4B-Instruct-2507
MedGemma-4B-IT
Medical
4 B
google/medgemma-4b-it
Llama-3.1-8B-Instruct
General
8 B
meta-llama/Llama-3.1-8B-Instruct
Bio-Medical-Llama-3-8B
Medical
8 B
ContactDoctor/Bio-Medical-Llama-3-8B
Appendix
Table 6: Models evaluated in the main sweep.
Figure 8 : Best tuning-window rewrite accuracy per (method × model).
Figure 9 : Pre-edit performance across task tiers, averaged over the five base models. Closed-form tasks are measured via exact-match accuracy, while open-form tasks (compositional, operational) are scored by an LLM judge and normalized to [0,1] .
Figure 10 : Generalization chain for LoRA-Merge, DPO, and GRPO.
Figure 11 : Per-tier post–pre performance over the weekly update stream. Bold lines show method-family means; faint lines show individual methods.
Figure 12 : Generalization results under daily and weekly batching. Each panel corresponds to one task tier.
Figure 13 : RAG failure-mode analysis on Llama-3.1-8B. The top row compares realistic RAG, empty-corpus RAG, Oracle-Llama, and Oracle-Opus; the bottom row reports retrieval accuracy under default and empty-corpus settings.
Tier
BM25
Dense
Anchor
0.40
0.27
Lexical
0.49
0.33
Relational
0.40
0.29
Compositional
0.32
0.35
Operational
0.00
0.03
Appendix
Table 7: Per-tier recall@ 3 on the realistic full corpus.
Figure 14 : Refusal rates by task and model across all evaluated task cases.
Figure 15 : CapTrack relative deviation (%) from the unedited base model, averaged across the five base models. Columns correspond to the full set of CAN / WILL / HOW capabilities together with per-category averages. Long-context tasks were not evaluated and therefore appear as 0 by construction.
Method
Model
Update
Lexical
Relational
Compositional
Operational
Parameter editing
AlphaEdit
Llama-3.1-8B
+9.1
−1.9
−7.4
−67.8
−28.3
Qwen-3-4B
−15.6
−19.3
−21.5
−72.6
−37.6
Gemma-3-4B
−16.0
−23.1
−15.6
−42.6
+0.6
MedGemma-4B
−29.1
−38.4
−37.0
−61.6
−5.1
Bio-Med-Llama-3-8B
−0.0
−1.8
−7.1
−49.6
−35.4
Appendix
Table 8 : Complete per-(method, model) generalization results underlying Figure 3 . Post–pre performance change (percentage points) per task tier; the two open-ended tiers (Compositional, Operational) are normalized to [0,1] before differencing. Oracle-retrieval rows are an upper-bound reference not drawn in Figure 3 .
Method
Model
Current Δ
Previous Δ
Drop
Parameter editing
AlphaEdit
Llama-3.1-8B
+0.8
−11.9
+12.8
Qwen-3-4B
−18.1
−21.5
+3.3
Gemma-3-4B
−18.9
−17.4
−1.5
MedGemma-4B
−34.5
−32.9
−1.6
Bio-Med-Llama-3-8B
−1.6
−9.9
+8.3
Appendix
Table 9 : Complete per-(method, model) sequential-retention results underlying Figure 4 . Closed-QA accuracy change (percentage points) vs. the per-model pre-edit baseline: Current Δ on the just-integrated update, Previous Δ on previously integrated updates (sentinel probe), and Drop = Current − Previous.
Method
Model
Update Δ
Locality Δ
Parameter editing
AlphaEdit
Llama-3.1-8B
+9.1
−47.7
Qwen-3-4B
−15.6
−39.4
Gemma-3-4B
−16.0
−20.1
MedGemma-4B
−29.1
−52.7
Bio-Med-Llama-3-8B
−0.0
−29.2
Appendix
Table 10 : Complete per-(method, model) update-integration and within-domain locality results underlying Figure 5 (a). Post–pre Δ (percentage points); a large positive Update with near-zero Locality indicates the target update was integrated without damaging neighboring oncology facts.
Method
Latent Competence
Behavioural Preferences
Protocol Compliance
Parameter editing
AlphaEdit
78.7
52.7
82.3
MEMIT
35.5
26.0
39.4
Augmented editing
GRACE
3.6
3.9
1.8
MEMOIR
4.6
3.5
5.0
Appendix
Table 11 : Complete out-of-domain capability forgetting results underlying Figure 5 (b). CapTrack relative deviation magnitude (%) vs. the unedited base model, averaged over models; larger is worse. Only degradations count (per-probe improvements are clipped to 0 before averaging), and weight-preserving methods (IKE, RAG variants) are 0 by construction.
Figure 16 : Deployment cost of knowledge integration methods under the weekly update setting. Each point corresponds to a method–model pair, showing average edit time per update (x-axis, log scale) against average inference time per evaluation case (y-axis, log scale). Two clear inference regimes emerge: methods compatible with vLLM achieve substantially lower deployment-time inference cost, while methods requiring custom HuggingFace forward passes incur one to two orders of magnitude slower generation. Retrieval methods exhibit negligible edit cost but remain constrained by inference-time retrieval and prompt length overheads.
The existing methods for evaluating the medical knowledge of Large Language Models (LLMs) are largely based on atemporal examination-style benchmarks, while in reality, medical knowledge is inherently dynamic and continuously evolves as new evidence emerges and treatments are approved. Consequently, evaluating medical knowledge without a temporal context may provide an incomplete assessment of whether LLMs can accurately reason about time-specific medical knowledge. Moreover, most medical data are historical, requiring the models not only to recall the correct knowledge, but also to know when that knowledge is correct. To bridge the gap, we built TempoMed-Bench, the first-of-its-kind benchmark for evaluating the temporal awareness of the LLMs in the medical domain through evolving guideline knowledge. Based on the TempoMed-Bench, our evaluation analysis first reveals that LLMs lack temporal awareness in medical knowledge through the key findings: (1) model performance on up-to-date medical knowledge exhibits a gradual linear decline over time rather than a sharp knowledge-cutoff behavior, suggesting that parametric medical knowledge is not strictly bounded by knowledge cutoffs; (2) LLMs consistently struggle more with recalling outdated historical medical knowledge than with up-to-date recommendations: accuracy of historical knowledge is only 25.37%-53.89% of up-to-date knowledge, indicating potential knowledge forgetting effects during training; and (3) LLMs often exhibit temporally inconsistent behaviors, where predictions fluctuate irregularly across neighboring years. We also show that the temporal awareness problem is a challenge that cannot be easily solved when integrated with agentic search tools (-3.15%-14.14%). This work highlights an important yet underexplored challenge and motivates future research on developing LLMs that can better encode time-specific medical knowledge.
Zihan Guan, Qiao Jin, Guangzhi Xiong +6
University of Virginia · National Institutes of Health · Dana-Farber Cancer Institute +2
Medical knowledge changes continually, making large language models vulnerable to relying on outdated yet clinically plausible information. We study whether the format of supervision affects medical knowledge updating under a matched training-budget setting. We introduce SEER-Bench, a temporally anchored oncology-staging benchmark curated from the latest versioned SEER Research Data release, and render identical medical update events from NCCN oncology guidelines into four supervision formats: EMQ, MSQ, FITB, and SAQ. Across SEER-Bench and HealthBench Professional, EMQ gives the most stable external transfer and retention among same-budget SFT variants. With EMQ supervision, the updated 4B model produces competitive results on temporally anchored oncology staging, reaching 64.8% answer accuracy and 59.6% rationale accuracy on SEER-Bench. Diagnostic analyses suggest that EMQ exposes denser clinical contrast signals while preserving discriminative representations with smaller movement from the base model. These results show that medical knowledge updating depends not only on the update algorithm, but also on how knowledge is structured as supervision.
Yangmin Huang, Shu Quan, He Geng +5
Xunfei Healthcare Technology Co., Ltd. · USTC-Xunfei Healthcare Digital and Health Joint Laboratory
A recent Nature Medicine study reports that general-purpose frontier LLMs outperform specialized retrieval-augmented clinical tools on medical benchmarks, and that retrieval can hurt strong models. We ask the natural follow-up: does structured knowledge-graph (KG) grounding change this, and when does grounding help at all? We contribute two results. First, a reproduction: the study's headline HealthBench score (~88) is the Consensus variant, not full HealthBench, where frontier models and ideal completions both score ~46-47 under a physician-calibrated grader (agreement 82.5%); we reproduce GPT-5.2 Consensus =90.9 and flag a score-deflating grader bug. Second, a knowledge-boundary result. Using a graph+vector engine (samyama-graph) over the public biomedical KG PrimeKG, neither naive triple retrieval nor an agentic natural-language-to-Cypher loop (82% successful queries) improves MedQA across a weak-to-strong model ladder (all |Delta| <= 3.4). On a synthetic counterfactual KG, and on a hybrid benchmark mixing known and novel facts, the identical pipeline lifts out-of-training accuracy from chance to ~100% (+68 to +79) while adding nothing on known facts (a no-LLM arm answers both). Across three regimes (no-knowledge, graph-aided, hybrid), grounding helps only insofar as the decisive fact lies outside the model's training -- public-KG facts are redundant, private and novel data are where it pays -- matching the study's institutional-data caveat.