MedKIT: Evaluating Knowledge Integration and Generalization in Large Language Models
Authors: Lukas Thede, Yash Kumar Atri, David Chen, Danielle Bitterman, Matthias Bethge, Tom Hartvigsen, Zeynep Akata
Organizations: University of Tübingen, Tübingen AI Center · Helmholtz Munich · Munich Center for Machine Learning (MCML) · University of Virginia · University of British Columbia · Harvard Medical School · Technical University of Munich
Constantly evolving real-world knowledge necessitates models to be updated continuously. Especially in medicine, as clinical evidence changes over time, outdated knowledge can pose safety risks. Existing evaluations of knowledge integration focus on factual recall, offering limited insight into whether newly integrated knowledge is actually usable. Our benchmark MedKIT (Medical Knowledge Integration and Transfer) provides a granular evaluation of how models integrate and apply knowledge under realistic sequences of clinical updates. Each instance corresponds to a factual update derived from clinical evidence, paired with targeted probes that assess transfer across lexical variation, relational transformations, compositional reasoning, and open-ended operationalization, as well as locality tests for knowledge preservation. Using MedKIT, we conduct a large-scale empirical study of 12 knowledge integration strategies across 5 diverse models, including both general-purpose and medical LLMs. Our results reveal a consistent gap between recall and usable knowledge: while most methods achieve strong gains on the original update task and under lexical variation, relational generalization is limited, and no method yields meaningful improvements on compositional or operational tasks. These findings highlight a fundamental challenge in knowledge integration and position MedKIT as a testbed for developing methods that make newly integrated knowledge more consistently usable across tasks and contexts.
Figures & tables
Figure 1 : MedKIT construction pipeline and examples. Structured clinical comparisons are converted into factual updates and multiple task variants probing lexical, relational, compositional, and operational generalization, alongside locality tests for knowledge preservation.
Benchmark
Domain
Input
Generalization
L
R
C
O
CounterFact [ 30 ]
G
Triples
✓
-
-
-
ZSRE [ 26 ]
G
QA
✓
-
-
-
WikiBigEdit [ 39 ]
G
QA / Wiki facts
✓
✓
✓
-
MQuAKE [ 54 ]
G
QA (multi-hop)
-
✓
-
-
DocTER [ 47 ]
G
Documents
-
-
-
-
Table 1 : Comparing factual knowledge integration benchmarks. Prior work emphasizes factual recall and isolated generalization settings, with limited coverage of multi-level transfer, open-ended tasks, and realistic update scenarios. (G: General, M: Medical, L: Lexical, R: Relational, C: Compositional, O: Operational)
Figure 2 : MedKIT statistics and temporal structure. The benchmark spans 6,196 clinically grounded factual updates across conditions and oncology groups, with balanced labels and temporally ordered batches that enable realistic sequential evaluation.
Figure 3 : Generalization of integrated knowledge across task formats. Mean post–pre performance ( Δ ) across tasks from direct recall ( Update ) to increasingly demanding generalization settings. Top: aggregated by method; bottom: aggregated by model, with faint markers indicating individual model–method combinations. Across both views, performance consistently degrades beyond lexical variation, indicating that improvements in recall do not translate into reliable generalization across reformulated and open-ended task settings.
Figure 4 : Retention under sequential updates. Current Δ measures immediate gains, Previous Δ performance on earlier updates. Retention varies by method, reflecting differences in underlying integration mechanisms.
Figure 5 : Capability preservation under knowledge integration. Left: Update success versus locality change, measuring whether methods integrate target updates while preserving neighboring oncology facts. Right: Changes in general capabilities, grouped into latent competence, behavioral preferences, and protocol compliance, highlight the risk of updates interfering with the model.
Method
Update
Lexical
Compos.
Locality
GRACE
+75.5
+0.0
+0.0
+0.0
MEMIT
+38.1
+22.8
0.0
-1.9
MEMOIR
+52.2
+38.1
-0.8
-2.2
SEEKR
+41.9
+20.2
-1.9
-0.9
Table 2: Cross-domain evaluation on WikiBigEdit. Mean post–pre change in exact-match containment (pp) across two models.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Strategy
# batches
Median size
Mean size
Max size
Daily
98
2
2.9
17
Weekly
48
4
5.8
21
Monthly
14
16.5
20.2
44
Appendix
Table 3: Batch-size statistics for the post-2025 evaluation window.
Judge
OQ κ ( k=3 )
OQ acc
OG κ ( k=2 )
OG acc
OG top-1
claude-opus-4
1.00
100 %
1.00
100 %
✓
claude-3-5-haiku
1.00
100 %
1.00
100 %
✓
gemini-2.0-flash
1.00
100 %
1.00
100 %
✓
gemini-2.5-pro
1.00
100 %
1.00
100 %
✓
gpt-4o-mini
1.00
100 %
1.00
100 %
✓
gpt-4o
1.00
100 %
0.42
86.7 %
✓
Appendix
Table 4: Cross-judge tier-classification agreement. Per-judge: rank-based LOJO Cohen’s κ and accuracy against the 6-judge consensus tier labels. The final column reports whether the judge identifies Oracle-RAG-abs as the strongest-performing OG method. Bottom row: Fleiss’ κ across all 7 judges’ independently-derived tier labels. k=3 on OQ (oracle / middle-performing / degradation), k=2 on OG (oracle / remaining methods).
Figure 6 : Consensus per-method judge scores (mean over 7 judges) with 95% tag-resampling bootstrap confidence intervals and tier bands. Left: compositional task (OQ, k=3 ). Right: operational task (OG, k=2 ). On OQ, oracle retrieval methods form a distinct top-performing tier, while AlphaEdit forms a separate degradation regime. On OG, the data support only a stable oracle-vs-rest separation, with the remaining methods clustering near pre-update performance.
Judge
OQ ρ
OQ κ
OQ MAE
OG ρ
OG κ
OG MAE
OG pref.
claude-opus-4
0.86
0.81
0.45
0.54
0.48
0.66
57.5%
gpt-4o
0.83
0.78
0.52
0.64
0.57
0.59
51.0%
claude-3-5-haiku
0.84
0.80
0.43
0.42
0.29
0.84
47.9%
llama-3.3-70b
0.84
0.80
0.46
0.45
0.34
0.92
53.0%
gemini-2.0-flash
0.77
0.73
0.54
0.54
0.48
0.66
52.5%
gemini-2.5-pro
0.77
0.79
0.57
0.64
0.52
0.72
59.5%
Appendix
Table 5: Agreement with the two-clinician consensus on the 100-item annotation set. For each judge we report Spearman correlation ( ρ ), quadratic-weighted Cohen’s κ , and mean absolute error (MAE) on the 1–5 scale against the mean of the two independent clinician annotators, plus operational-task agreement on the categorical preference flag (OG pref.), averaged over the two clinicians. The final row reports inter-clinician agreement as a human agreement reference.
Figure 7 : Pairwise Spearman ρ between the seven candidate judges and two independent clinicians on the 100-item annotation set. Left: compositional OQ. Right: operational OG. Inter-clinician agreement is ρ=0.76 on OQ and ρ=0.62 on OG; judge agreement is generally higher on the structured compositional task than on open-ended operational generation.
Model
Domain
Scale
HuggingFace ID
Gemma-3-4B-IT
General
4 B
google/gemma-3-4b-it
Qwen-3-4B-Instruct
General
4 B
Qwen/Qwen3-4B-Instruct-2507
MedGemma-4B-IT
Medical
4 B
google/medgemma-4b-it
Llama-3.1-8B-Instruct
General
8 B
meta-llama/Llama-3.1-8B-Instruct
Bio-Medical-Llama-3-8B
Medical
8 B
ContactDoctor/Bio-Medical-Llama-3-8B
Appendix
Table 6: Models evaluated in the main sweep.
Figure 8 : Best tuning-window rewrite accuracy per (method × model).
Figure 9 : Pre-edit performance across task tiers, averaged over the five base models. Closed-form tasks are measured via exact-match accuracy, while open-form tasks (compositional, operational) are scored by an LLM judge and normalized to [0,1] .
Figure 10 : Generalization chain for LoRA-Merge, DPO, and GRPO.
Figure 11 : Per-tier post–pre performance over the weekly update stream. Bold lines show method-family means; faint lines show individual methods.
Figure 12 : Generalization results under daily and weekly batching. Each panel corresponds to one task tier.
Figure 13 : RAG failure-mode analysis on Llama-3.1-8B. The top row compares realistic RAG, empty-corpus RAG, Oracle-Llama, and Oracle-Opus; the bottom row reports retrieval accuracy under default and empty-corpus settings.
Tier
BM25
Dense
Anchor
0.40
0.27
Lexical
0.49
0.33
Relational
0.40
0.29
Compositional
0.32
0.35
Operational
0.00
0.03
Appendix
Table 7: Per-tier recall@ 3 on the realistic full corpus.
Figure 14 : Refusal rates by task and model across all evaluated task cases.
Figure 15 : CapTrack relative deviation (%) from the unedited base model, averaged across the five base models. Columns correspond to the full set of CAN / WILL / HOW capabilities together with per-category averages. Long-context tasks were not evaluated and therefore appear as 0 by construction.
Method
Model
Update
Lexical
Relational
Compositional
Operational
Parameter editing
AlphaEdit
Llama-3.1-8B
+9.1
−1.9
−7.4
−67.8
−28.3
Qwen-3-4B
−15.6
−19.3
−21.5
−72.6
−37.6
Gemma-3-4B
−16.0
−23.1
−15.6
−42.6
+0.6
MedGemma-4B
−29.1
−38.4
−37.0
−61.6
−5.1
Bio-Med-Llama-3-8B
−0.0
−1.8
−7.1
−49.6
−35.4
Appendix
Table 8 : Complete per-(method, model) generalization results underlying Figure 3 . Post–pre performance change (percentage points) per task tier; the two open-ended tiers (Compositional, Operational) are normalized to [0,1] before differencing. Oracle-retrieval rows are an upper-bound reference not drawn in Figure 3 .
Method
Model
Current Δ
Previous Δ
Drop
Parameter editing
AlphaEdit
Llama-3.1-8B
+0.8
−11.9
+12.8
Qwen-3-4B
−18.1
−21.5
+3.3
Gemma-3-4B
−18.9
−17.4
−1.5
MedGemma-4B
−34.5
−32.9
−1.6
Bio-Med-Llama-3-8B
−1.6
−9.9
+8.3
Appendix
Table 9 : Complete per-(method, model) sequential-retention results underlying Figure 4 . Closed-QA accuracy change (percentage points) vs. the per-model pre-edit baseline: Current Δ on the just-integrated update, Previous Δ on previously integrated updates (sentinel probe), and Drop = Current − Previous.
Method
Model
Update Δ
Locality Δ
Parameter editing
AlphaEdit
Llama-3.1-8B
+9.1
−47.7
Qwen-3-4B
−15.6
−39.4
Gemma-3-4B
−16.0
−20.1
MedGemma-4B
−29.1
−52.7
Bio-Med-Llama-3-8B
−0.0
−29.2
Appendix
Table 10 : Complete per-(method, model) update-integration and within-domain locality results underlying Figure 5 (a). Post–pre Δ (percentage points); a large positive Update with near-zero Locality indicates the target update was integrated without damaging neighboring oncology facts.
Method
Latent Competence
Behavioural Preferences
Protocol Compliance
Parameter editing
AlphaEdit
78.7
52.7
82.3
MEMIT
35.5
26.0
39.4
Augmented editing
GRACE
3.6
3.9
1.8
MEMOIR
4.6
3.5
5.0
Appendix
Table 11 : Complete out-of-domain capability forgetting results underlying Figure 5 (b). CapTrack relative deviation magnitude (%) vs. the unedited base model, averaged over models; larger is worse. Only degradations count (per-probe improvements are clipped to 0 before averaging), and weight-preserving methods (IKE, RAG variants) are 0 by construction.
Figure 16 : Deployment cost of knowledge integration methods under the weekly update setting. Each point corresponds to a method–model pair, showing average edit time per update (x-axis, log scale) against average inference time per evaluation case (y-axis, log scale). Two clear inference regimes emerge: methods compatible with vLLM achieve substantially lower deployment-time inference cost, while methods requiring custom HuggingFace forward passes incur one to two orders of magnitude slower generation. Retrieval methods exhibit negligible edit cost but remain constrained by inference-time retrieval and prompt length overheads.