Clinical Note Bloat Reduction for Efficient LLM Use
Authors: Jordan L. Cahoon, Chloe Stanwyck, Asad Aali, Rachel Madding, Sulaiman S. Somani, Emma Sun, Yixing Jiang, Renumathy Dhanasekaran, +1 more
Organizations: Department of Biomedical Data Science, Stanford University, Stanford, CA · Department of Pathology, Stanford University, Stanford, CA · Department of Anesthesiology, Perioperative and Pain Medicine, Stanford University, Stanford, CA · Department of Radiology, Stanford University, Stanford, CA · Department of Obstetrics and Gynecology, Stanford University, Stanford, CA · Department of Computer Science, Stanford University, Stanford, CA · Division of Gastroenterology and Hepatology, Stanford University, Stanford, CA · Weill Cancer Hub West
Background: Clinical notes contain extensive duplicated text from templates, copy-paste, and auto-populated fields ("note bloat"), diluting clinical signal, limiting longitudinal context, and increasing large language model (LLM) costs. Methods: TRACE removes note bloat using note-level EHR metadata to identify templated and copied content, with frequency-based de-duplication when metadata are unavailable. We evaluated TRACE using blinded physician span review and gold-standard templated-text annotations across four cohorts spanning liver transplant, obstetrics, and inpatient populations at multiple health systems (5.3M notes). We compared zero-shot LLMs and embedding-based classifiers using original and TRACE-processed notes for 20 information extraction tasks and prediction of 5-year survival, postpartum hemorrhage, and 30-day readmission. Results: Only 0.3-6.6% of removed text was flagged as author-generated; TRACE captured 86% of annotated templated characters. Information extraction F1 differences averaged by cohort ranged from -0.009 to +0.004; task-specific prediction F1 differences ranged from -0.011 to +0.018. Among 1,000 randomly sampled Stanford Health Care patients, TRACE reduced chart text by 47.3% (742.7M characters), averaging 220,167 fewer tokens per patient. Using 2024 encounter volumes at a large tertiary academic center and one query per encounter, projected three-year net savings ranged from 1.00Mto13.58M across evaluated model pricing schemes, including initial and annual TRACE processing costs. Conclusion: TRACE substantially reduces clinical note redundancy while preserving information extraction and prediction performance. Underused EHR metadata can reduce LLM inference costs, expand usable longitudinal context, and support scalable clinical AI.
Figures & tables
Figure 1: TRACE Method. (A) TRACE identifies templated and copied text using Epic Clarity attribution metadata (Reference Module) and flags additional redundant text blocks using frequency-based de-duplication (Frequency Module). (B) Example TRACE-annotated note showing templated and copied spans with sources identified by the Reference and Frequency Modules.
Figure 2: TRACE downstream tasks. (A) TRACE utility is evaluated through two downstream tasks: Information Extraction (IE) and Clinical Outcome Prediction. In IE, a large language model performs inference over batched notes and updates the prediction as new notes are added. In clinical outcome prediction, the most recent 32K tokens (embedding-based) or up to 1M tokens (zero-shot) are used. Notes can either be used for zero-shot inference or embedded and used to train a classifier. (B) Datasets for downstream evaluation.
Figure 3: TRACE facilitates the analysis of templated and copied text in 1,000 randomly sampled patients (A) Distribution of text removal by TRACE on full note timelines. (B) Number of characters identified as Templated, Copied, and the intersection. (C) Number of characters identified with Reference and Frequency Modules. Counts reflect flagged characters, while the intersection refers to characters identified by both methods. Because copied spans are retained at their first occurrence, flagged characters exceed removed characters. (D) Percent of templated and copied text for the 24 most common note types in Stanford Health Care.
Figure 4: TRACE-processed notes achieve non-inferior performance in information extraction (IE) compared with original notes (A) IE performance across tasks in hepatology and obstetrics for original and TRACE Notes. The X axis indicates the F1 score for the respective cohort. The Y axis indicates the specific task. (B) Distribution of text removed in patient timelines across four real world cohorts. (C) Paired F1 differences (TRACE – Original), averaged across IE tasks, by within-cohort quartile of the percentage of characters removed by TRACE (Q1, least; Q4, most). Points show observed differences; error bars show 95% percentile bootstrap confidence intervals (1,000 replicates).
Figure 5: TRACE enables efficient clinical outcome prediction. A) ROC curves for Postpartum Hemorrhage and 30-day Readmission B) F1 scores with 95% bootstrapped confidence intervals for zero-shot Overall Survival prediction (*) and embedding-based PPH and readmission prediction C) Paired F1 differences (TRACE minus Original) by within-cohort quartile of the percentage of characters removed by TRACE (Q1, least; Q4, most). Points show observed differences; error bars show 95% percentile bootstrap confidence intervals (1,000 replicates).
Supplementary Figure 1: Example of TRACE span identification. (A) Source template. (B) Corresponding de-identified note with templated, copy-forward, and overlapping spans highlighted. Edits have been made to ensure de-identification, changing some wording and minor information while retaining identified spans. Copy-forward spans are retained at their first occurrence. Highlights show identified spans, while removal depends on module-specific length thresholds.
Supplementary Figure 2: Prompt used with GPT-4.1 to draft candidate templated and structured-data spans for the metadata-assisted reference annotations.
Supplementary Figure 3: Prompt used for information extraction for the first batch
Supplementary Figure 4: Prompt used to update extracted values after the first batch
Supplementary Figure 5: Prompt used for zero-shot prediction of outcomes.
Written by
Method
Structured
Template
note author
Indeterminate
Annotator 1
TRACE
Reference
0.06
91.27
1.96
6.71
(0.00–0.18)
(85.58–95.69)
(0.57–4.06)
(2.87–11.82)
Frequency
3.08
90.11
0.29
6.52
(1.05–6.13)
(82.93–94.70)
(0.00–0.93)
(2.86–12.07)
Supplementary Table 1: Annotation category composition for TRACE (Reference and Frequency selection) versus Random selection. Values are percentages of labeled characters (95% confidence intervals), reported separately by annotator. The highest percentage in each row is shown in bold.
Minimum span (characters)
Method
Precision (95% CI)
Recall (95% CI)
≥ 25
Reference
0.9568 (0.9303, 0.9747)
0.8580 (0.7725, 0.9152)
≥ 25
TRACE
0.8789 (0.8139, 0.9254)
0.8763 (0.7967, 0.9299)
≥ 50*
Reference
0.9614 (0.9343, 0.9792)
0.8423 (0.7546, 0.9024)
≥ 50*
TRACE
0.8770 (0.8055, 0.9265)
0.8643 (0.7816, 0.9207)
≥ 75
Reference
0.9623 (0.9330, 0.9812)
0.7675 (0.6645, 0.8429)
≥ 75
TRACE
0.8930 (0.8228, 0.9395)
0.7946 (0.6927, 0.8663)
Supplementary Table 2: Template-detection precision and recall across minimum span thresholds in the 100-note annotated subset. Reference, Reference Module only; TRACE, Reference and Frequency Modules. Recall was evaluated against the same annotated templated spans ( ≥ 50 characters) at every threshold. TRACE precision is conservative because templated text flagged only by the Frequency Module is not attribution-linked and therefore counts as a false positive. Values are estimates (95% confidence intervals from 10,000 note-level bootstrap replicates).
Minimum span (characters)
Corpus removal (95% CI)
Timeline removal (95% CI)
≥ 25
48.97% (46.53%, 51.45%)
40.58% (39.73%, 41.42%)
≥ 50*
47.25% (44.77%, 49.76%)
38.85% (38.01%, 39.67%)
≥ 75
45.06% (42.60%, 47.57%)
36.44% (35.59%, 37.27%)
Supplementary Table 3: Text removal across minimum span thresholds in the 1,000-patient task-agnostic sample. Corpus removal is the percentage of all characters removed; timeline removal is the mean per-patient percentage removed. Values are estimates (95% confidence intervals from 10,000 patient-level bootstrap replicates).
type
mean
std
min
25
50
75
max
Progress Notes (n=120467)
0.51
0.3
0.0031
0.23
0.5
0.8
1.0
Telephone Encounter (n=34780)
0.36
0.31
0.0013
0.091
0.25
0.6
1.0
Patient Instructions (n=15399)
0.72
0.28
0.0078
0.53
0.82
0.95
1.0
Consults (n=12202)
0.29
0.24
0.0039
0.1
0.21
0.4
1.0
H&P (n=4633)
0.31
0.24
0.0039
0.13
0.25
0.44
1.0
ED Provider Notes (n=4156)
0.2
0.15
0.0042
0.092
0.16
0.26
0.96
Supplementary Table 4: Annotation by TRACE in 1000 patients for 30 most frequent note types
type
mean
std
min
25
50
75
max
Progress Notes (n=120467)
0.16
0.22
0.0
0.023
0.075
0.2
1.0
Telephone Encounter (n=34780)
0.27
0.29
0.0
0.061
0.13
0.4
1.0
Patient Instructions (n=15399)
0.62
0.33
0.0
0.35
0.72
0.91
1.0
Consults (n=12202)
0.14
0.15
0.0
0.045
0.094
0.18
1.0
H&P (n=4633)
0.098
0.13
0.0
0.028
0.063
0.12
1.0
ED Provider Notes (n=4156)
0.1
0.079
0.0
0.054
0.083
0.13
0.91
Supplementary Table 5: Template annotation by TRACE in 1000 patients for 30 most frequent note types
type
mean
std
min
25
50
75
max
Progress Notes (n=120467)
0.37
0.36
0.0
0.0
0.26
0.73
1.0
Telephone Encounter (n=34780)
0.11
0.26
0.0
0.0
0.0
0.0
1.0
Patient Instructions (n=15399)
0.13
0.3
0.0
0.0
0.0
0.0
1.0
Consults (n=12202)
0.16
0.24
0.0
0.0
0.031
0.22
1.0
H&P (n=4633)
0.22
0.25
0.0
0.021
0.14
0.35
1.0
ED Provider Notes (n=4156)
0.095
0.14
0.0
0.0
0.031
0.13
0.96
Supplementary Table 6: Copyforward annotation by TRACE in 1000 patients for 30 most frequent note types
Cohort
Metric
Δ
95% CI for Δ
OB/PPH
F1
0.0180
( 0.0023 , 0.0341 )
OB/PPH
AUROC
0.0047
( −0.0034 , 0.0127 )
SHC-Inpatient
F1
−0.0020
( −0.0141 , 0.0103 )
SHC-Inpatient
AUROC
0.0003
( −0.0052 , 0.0059 )
MIMIC
F1
0.0070
( −0.0106 , 0.0244 )
MIMIC
AUROC
−0.0033
( −0.0147 , 0.0080 )
Supplementary Table 7: Outcome-prediction performance differences with 95% confidence intervals.
Cohort
Task
Δ
95% CI for Δ
Transplant
Average IE
0.0040
( −0.0047 , 0.0126 )
Transplant
alcoholic fatty liver disease
0.0210
( −0.0341 , 0.0784 )
Transplant
ascites
0.0094
( −0.0100 , 0.0308 )
Transplant
cirrhosis
−0.0037
( −0.0093 , 0 )
Transplant
cryptogenic cirrhosis
0
( 0 , 0 )
Transplant
hepatic encephalopathy
0.0091
( 0 , 0.0241 )
Supplementary Table 8: Aggregate and individual information-extraction F1 differences with 95% confidence intervals.
Supplementary Figure 6: TRACE 95% confidence interval for Information Extraction F1 compared to distribution of F1 for random removal
Term
β (SE)
OR (95% CI)
p
pHolm
Family 1: Information extraction (2 hypotheses)
Transplant: average IE
Intercept
0.863 ( 0.1565 )
2.37 ( 1.74 , 3.22 )
3.529×10−8
—
TRACE
0.0496 ( 0.0394 )
1.05 ( 0.973 , 1.14 )
0.209
0.4179
Prevalence
2.1238 ( 0.2854 )
8.36 ( 4.78 , 14.6 )
9.869×10−14
—
Obstetrics: average IE
Supplementary Table 9: Logistic regression coefficients for TRACE versus Original. Endpoint: sensitivity. IE denotes information extraction; prevalence is included only in IE models.
Term
β (SE)
OR (95% CI)
p
pHolm
Family 1: Information extraction (2 hypotheses)
Transplant: average IE
Intercept
−1.0865 ( 0.2572 )
0.337 ( 0.204 , 0.558 )
2.392×10−5
—
TRACE
0.0048 ( 0.0291 )
1 ( 0.949 , 1.06 )
0.8696
0.8696
Prevalence
4.6439 ( 0.3721 )
104 ( 50.1 , 216 )
9.442×10−36
—
Obstetrics: average IE
Supplementary Table 10: Logistic regression coefficients for TRACE versus Original. Endpoint: specificity. IE denotes information extraction; prevalence is included only in IE models.
Cohort / task
df
Wald χ2
p
pHolm
Family 1: Information extraction (2 hypotheses)
Transplant: average IE
4
4.1295
0.3888
0.3888
Obstetrics: average IE
4
9.5479
0.04877
0.09755
Family 2: Zero-shot outcome prediction (1 hypothesis)
Transplant: survival*
1
3.2908
0.06967
0.06967
Family 3: Embedding-based outcome prediction (3 hypotheses)
Supplementary Table 11: Wald tests of text-removal associations among TRACE predictions. Endpoint: sensitivity. *Linear term for text removal (1 df) because matrix was singular.
Cohort / task
df
Wald χ2
p
pHolm
Family 1: Information extraction (2 hypotheses)
Transplant: average IE
4
3.1419
0.5344
1
Obstetrics: average IE
4
0.7973
0.9388
1
Family 2: Zero-shot outcome prediction (1 hypothesis)
Transplant: survival
4
2.506
0.6436
0.6436
Family 3: Embedding-based outcome prediction (3 hypotheses)
Supplementary Table 12: Wald tests of text-removal associations among TRACE predictions. Endpoint: specificity.
Supplementary Figure 7: Association between text removal and information extraction accuracy. Adjusted predicted probability of a correct prediction among TRACE-processed inputs as a function of the percentage of characters removed by TRACE, estimated from restricted cubic spline logistic regression models (Equation 3) fit separately among reference-positive (sensitivity) and reference-negative (specificity) observations. Original input length was held at its mean and feature prevalence at the task-specific prevalence. Thick lines show the average across IE tasks; thin lines show individual tasks. Shading denotes pointwise 95% confidence intervals from the patient-cluster-robust covariance matrix.
Supplementary Figure 8: Association between text removal and clinical outcome prediction accuracy. Adjusted predicted probability of a correct prediction among TRACE-processed inputs as a function of the percentage of characters removed by TRACE, estimated from restricted cubic spline logistic regression models (Equation 3) fit separately among reference-positive (sensitivity) and reference-negative (specificity) observations, with original input length held at its mean. Shading denotes pointwise 95% confidence intervals from the patient-cluster-robust covariance matrix. Overall survival sensitivity uses a linear term for text removal because the spline model was singular.
Cohort / task
Metric
Δ
Lower 97.5% bound
p
pHolm
Non- inferior
Family 1: Information extraction F1 (2 hypotheses)
Transplant: average IE
F1
0.0040
−0.0046
4.882×10−35
9.764×10−35
Yes
Obstetrics: average IE
F1
−0.0086
−0.0222
1.277×10−9
1.277×10−9
Yes
Family 2: Outcome-prediction AUROC (3 hypotheses)
Obstetrics: PPH
AUROC
0.0047
−0.0033
6.451×10−41
1.290×10−40
Yes
Inpatient-SHC: readmission
AUROC
0.0003
−0.0053
6.640×10−71
1.992×10−70
Yes
Supplementary Table 13: Exploratory Wald non-inferiority tests of TRACE versus Original at an absolute margin of 0.05. Δ denotes TRACE minus Original performance. For each endpoint, we tested H0:Δ≤−0.05 against H1:Δ>−0.05 . Holm correction was applied separately within each family shown; non-inferiority was concluded when the adjusted one-sided p -value was below 0.025. Family-wise error was controlled within each family, rather than jointly across all nine hypotheses. *Survival prediction used zero-shot inference; the other outcome-prediction tasks used embedding-based classifiers.
feature
arm
precision
precision_lower
precision_upper
Overall Survival (Liver Transplant)
Original
0.22
0.17
0.28
Overall Survival (Liver Transplant)
TRACE
0.22
0.17
0.27
Postpartum Hemorrhage (Obstetrics)
Original
0.30
0.28
0.32
Postpartum Hemorrhage (Obstetrics)
TRACE
0.34
0.32
0.37
30-Day Readmission (Inpatient-SHC)
Original
0.52
0.49
0.54
30-Day Readmission (Inpatient-SHC)
TRACE
0.48
0.46
0.5
Supplementary Table 14: Clinical outcome prediction precision with 95% confidence intervals. OS prediction was evaluated using zero-shot inference.
feature
arm
recall
recall_lower
recall_upper
Overall Survival (Liver Transplant)
Original
0.94
0.87
1.000
Overall Survival (Liver Transplant)
TRACE
0.91
0.82
0.98
Postpartum Hemorrhage (Obstetrics)
Original
0.59
0.56
0.62
Postpartum Hemorrhage (Obstetrics)
TRACE
0.53
0.50
0.56
30-Day Readmission (Inpatient-SHC)
Original
0.63
0.61
0.66
30-Day Readmission (Inpatient-SHC)
TRACE
0.68
0.66
0.7
Supplementary Table 15: Clinical outcome prediction recall with 95% confidence intervals. OS prediction was evaluated using zero-shot inference.
feature
arm
f1
f1_lower
f1_upper
Overall Survival (Liver Transplant)
Original
0.36
0.29
0.43
Overall Survival (Liver Transplant)
TRACE
0.35
0.28
0.42
Postpartum Hemorrhage (Obstetrics)
Original
0.40
0.38
0.42
Postpartum Hemorrhage (Obstetrics)
TRACE
0.42
0.39
0.44
30-Day Readmission (Inpatient-SHC)
Original
0.57
0.55
0.59
30-Day Readmission (Inpatient-SHC)
TRACE
0.57
0.55
0.58
Supplementary Table 16: Clinical outcome prediction F1 with 95% confidence intervals. OS prediction was evaluated using zero-shot inference.
feature
arm
auroc
auroc_lower
auroc_upper
Postpartum Hemorrhage (Obstetrics)
Original
0.747
0.729
0.763
Postpartum Hemorrhage (Obstetrics)
TRACE
0.751
0.734
0.768
30-Day Readmission (Inpatient-SHC)
Original
0.802
0.791
0.814
30-Day Readmission (Inpatient-SHC)
TRACE
0.803
0.791
0.814
30-Day Readmission (MIMIC-III)
Original
0.712
0.687
0.735
30-Day Readmission (MIMIC-III)
TRACE
0.708
0.685
0.731
Supplementary Table 17: Clinical outcome prediction AUROC with two-sided 95% percentile-bootstrap confidence intervals from 10,000 paired resampling replicates.
Supplementary Table 19: Embedding-based postpartum hemorrhage (PPH) prediction across TRACE minimum span thresholds. Values are estimates (95% confidence intervals). Original denotes unprocessed notes.
Supplementary Figure 9: Estimated net system-wide cost savings with TRACE. Gross annual savings were estimated from the mean reduction of 220,167 tokens per patient timeline (based on the 1,000-patient timeline sample), 2,058,497 FY2024 Stanford Health Care encounters (from published utilization data), and the model prices and query frequencies shown. Each query was assumed to process a full timeline. Processing 435,797 notes from 1,000 patients required 130.746 billable vCPU-hours. Assuming 8 GiB per vCPU gives 1,045.968 GiB-hours and a sample processing cost of 4.47atGoogleCloudE2virtualmachine1−yearcommitted−useratesunderanexistinginstitutionalcommitment.Extrapolationto3millionSHCpatientsgives1.307billionbaselinenotesandaninitialcostof13,403. New notes equal to a fixed 9% of baseline volume (approximately 117 million notes) adds a fixed cost of $1,206 annually, assuming stable model pricing. Net savings subtract initial and annual processing costs beginning in year 1. Projections assume representative sample timelines, constant average processing cost per note, incremental processing without rerunning the historical corpus, and constant query volume and token reduction. Storage and other operational costs are excluded.
School of Computer Science and Engineering, Bangor University, Bangor, LL57 2DG, Gwynedd, United Kingdom. · University of Thi-Qar, Nasiriyah, 64001, Iraq.