Auditable Clinical Timeline Reconstruction with Provenance-Aware Evidence Graphs
Organizations: Public Health Unit, Medical Information Department, Bordeaux University Hospital (CHU Bordeaux), Bordeaux, France.
Abstract
A patient-timeline reconstruction system is auditable only if it keeps the mentions behind each answer, records how facts were revised, and declines to answer when the evidence is not in the text. This study tests these three properties on a fully synthetic corpus (1,000 patients, 3,353 notes, 220 revision edges). Two provenance-aware Evidence Graph operators reduced the node-plus-edge count to 67% and 63% (77-78% of serialized size) while preserving every answer and mention link across 6,813 query points answerable by recency; a fixed-window baseline returned no value for 53.4% of points, unflagged. On evidence-unavailable controls that announce the omission, a BioClinicalBERT gate and a zero-shot LLM gate responded mainly to the announcement. On marker-free controls, BERT abstained on 0 of 81 notes while its accuracy fell from 93.8% to 59.3% across all three relation classes; the LLM's coverage fell from 75.6% to 27.7% on notes its own model family judged undeterminable. Against 482 regenerated gold spans, the LLM's cited evidence reached recall 0.850 and precision 0.864; BERT's span head, trained without span labels, did not localize evidence. A temporally versioned provenance graph stored abstentions as typed, queryable edges. The clean task admits a 0.651-accuracy shortcut, and results describe implementation behaviour on synthetic data, not clinical performance.
Figures & tables
| Quantity | Value |
|---|---|
| Patients | 1,000 |
| Notes (total) | 3,353 (realized mean 3.35 per patient) |
| Clinical facts | 2,220 |
| Patients with a documented revision | 110 (11.0%) |
| Revision edges (ground truth) | 220 |
| Relation-eligible notes (gold relation and span) | 1,000 |
| A. All query points ( ) | |||||||
| Representation | Nodes | Edges | Total | Size ratio | Query acc. | Query cov. | Prov. recall |
| Uncompressed | 6,926 | 7,146 | 14,072 | 1.00 | 1.000 | 1.000 | 1.000 |
| Aggregation | 2,220 | 7,146 | 9,366 | 0.67 | 1.000 | 1.000 | 1.000 |
| Revision-representation | 2,000 | 6,926 | 8,926 | 0.63 | 1.000 | 1.000 | 1.000 |
| Temporal-approximation | 2,000 | 6,978 | 8,978 | 0.64 | 0.987 | 0.996 | 1.000 |
| Fixed-window ( ) | 3,982 | 4,202 | 8,184 | 0.58 | 0.453 | 1.000 | 0.575 |
| System | Configuration | Abstain rate | Accuracy (answered) | Evidence recall | |
|---|---|---|---|---|---|
| BERT | Baseline (ungated) | 150 | 0.000 | 0.740 | 0.000 |
| BERT | Provenance-first (hard-gated) | 150 | 0.960 | 1.000 (6 answered) | 0.000 |
| LLM | Baseline (ungated) | 60 | 0.000 | 0.617 | 0.000 |
| LLM | Provenance-first (gated) | 60 | 0.550 | 0.556 (27 answered) | 0.000 |
| Gate | Rows | scorable | Recall | Precision | Span / note length |
|---|---|---|---|---|---|
| BERT | Ungated baseline | 77 / 150 | 0.049 | — | — |
| BERT | Availability-supervised (gated) | 77 / 150 | 0.006 | — | — |
| LLM | All scorable (abstentions scored on their cited span) | 482 / 1,000 | 0.750 | 0.866 | 0.329 |
| LLM | Answered rows only | 379 / 1,000 | 0.850 | 0.864 | 0.339 |
| System | Gate | Prompt | Coverage | False abst. | Answered acc. ( ans.) | |
|---|---|---|---|---|---|---|
| BERT | Ungated baseline | — | 150 | 1.000 | 0.000 | 0.740 (150) |
| BERT | Hard-gated | — | 150 | 0.040 | 0.960 | 1.000 (6) |
| BERT | Availability-supervised | — | 150 | 0.987 [0.953, 0.996] | 0.013 | 0.899 [0.840, 0.938] (148) |
| LLM | Ungated baseline | original | 60 | 1.000 | 0.000 | 0.617 (60) |
| LLM | Original gate | original | 60 | 0.450 | 0.550 | 0.556 (27) |
| LLM | Availability-aware | no event defs. | 1,000 | 0.587 [0.556, 0.617] | 0.413 | 0.373 [0.335, 0.413] (587) |
| Gate | Condition | Clean cov. | Abstention recall | Unsafe-answer rate | ans. | Answered accuracy | |
|---|---|---|---|---|---|---|---|
| BERT | remove_evidence | 150 | 0.987 | 1.000 | 0.000 | 0 | undefined |
| BERT | mask_temporal_cue | 150 | 0.987 | 0.980 [0.943, 0.993] | 0.007 [0.001, 0.037] | 3 | 0.667 |
| BERT | token inserted | 61 | 0.967 | 1.000 [0.941, 1.000] | 0.000 [0.000, 0.059] | 0 | undefined |
| BERT | notice, note intact | 89 | 1.000 | 0.966 [0.906, 0.988] | 0.011 [0.002, 0.061] | 3 | 0.667 |
| BERT | All controls | 300 | 0.987 | 0.990 [0.971, 0.997] | 0.003 [0.001, 0.019] | 3 | 0.667 |
| LLM | remove_evidence | 1,000 | 0.700 | 1.000 | 0.000 | 0 | undefined |
| Gate | Condition | Clean cov. | Clean acc. | Coverage (silent) | Accuracy (silent) | |
|---|---|---|---|---|---|---|
| BERT | silent_keep | 70 | 0.986 | 0.928 | 0.957 [0.881, 0.985] | 0.836 [0.729, 0.906] |
| BERT | silent_removal , implied_only | 44 | 0.977 | 0.930 | 1.000 [0.920, 1.000] | 0.614 [0.466, 0.743] |
| BERT | silent_removal , strict | 37 | 1.000 | 0.946 | 1.000 [0.906, 1.000] | 0.568 [0.409, 0.713] |
| BERT | silent_removal , all | 81 | 0.988 | 0.938 | 1.000 [0.955, 1.000] | 0.593 [0.484, 0.693] |
| LLM | silent_keep | 452 | 0.723 | 0.881 | 0.721 [0.678, 0.761] | 0.850 [0.807, 0.884] |
| LLM | silent_removal , implied_only | 269 | 0.651 | 0.851 | 0.688 [0.630, 0.740] | 0.384 [0.317, 0.456] |
| silent_removal | silent_keep | ||||||
| Gold | Clean | Edited acc. | Answered as | Clean | Edited | ||
| relation | acc. | [95% CI] | E1 E2/E2 E1/Ovl. | acc. | acc. | ||
| E1 E2 | 28 | 1.000 | 0.750 [0.566, 0.873] | 21/0/7 | 26 | 1.000 | 0.960 |
| Ovl. | 23 | 0.913 | 0.609 [0.408, 0.778] | 9/0/14 | 19 | 0.842 | 0.882 |
| E2 E1 | 30 | 0.900 | 0.433 [0.274, 0.608] | 4/13/13 | 25 | 0.920 | 0.680 |
| All | 81 | 0.938 | 0.593 [0.484, 0.693] | 34/13/34 | 70 | 0.928 | 0.836 |
| Quantity | Value |
|---|---|
| Total edges | 120 |
| Abstained edges | 33 (27.5%) |
| Relation types among non-abstained edges | EVENT1-BEFORE-EVENT2: 36; EVENT1-OVERLAP-EVENT2: 31; EVENT2-BEFORE-EVENT1: 20 |
| Mean coverage (character-span fraction; not comparable to BERT token-mask coverage) | 0.317 |
| Disagreements flagged (baseline versus gated) | 33 (all abstention-driven) |
| Metric | BERT baseline | BERT gated | LLM baseline | LLM gated |
|---|---|---|---|---|
| Certification examples | 30 | 30 | 20 | 20 |
| Certification | 0.01 | 0.01 | 0.05 | 0.05 |
| Abstained before perturbation | 0/30 | 28/30 | 0/20 | 16/20 |
| Certification rate | 0.000 | 0.000 | 0.250 | 0.000 |
| Mean prediction confidence | 0.926 | 0.042 | 0.939 | 0.158 |
| Mean provenance confidence | 0.567 | 0.033 | 0.850 | 0.158 |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Script | Role |
|---|---|
| sythetic_data_generation.py | Corpus generation (Section 2.2) |
| provenance_first_Bio_Clinical_BERT.py | Hard-gated BERT extraction, certification (Sections 2.6, 2.10) |
| provenance_first_LLM.py | Gated LLM extraction, certification (Sections 2.6, 2.10) |
| compression_operators.py | Evidence Graph compression operators and evaluation, including serialized size and revision subsets (Section 2.7) |
| provenance_temporal_kg.py | Provenance graph, entity resolution, contradiction flags, federation-ready aggregation (Section 2.9) |
| make_evidence_unavailable_adversarial.py | Evidence-unavailable (marker-carrying) control generation (Section 2.3) |