PrivMeSA: Privacy-Aware Self-Evolving Multi-Agent System for Medicine via Local-Remote LLM Collaboration
Authors: Dannong Wang, Yuran Zhang, Bian Sun, Alex Stinard, Yuzhang Shang, Song Wang, Yu Tian
Organizations: Institute of Artificial Intelligence, University of Central Florida · Department of Computer Science and Operations Research, Universit´e de Montr´eal · Department of Medicine, University of Central Florida
Clinical large language model (LLM) agents deployed locally can consult more capable remote models, but doing so risks exposing patient information. Privacy-conscious delegation places disclosure decisions with a local agent, yet removing explicit identifiers is insufficient: quasi-identifiers can accumulate across multi-turn consultations and repeated patient visits to enable re-identification. We introduce PrivMeSA, a privacy-aware self-evolving multi-agent system that learns to control disclosure and retains remote expertise for local reuse. A local agent manages each encounter and consults remote specialists that may request additional information. Reinforcement learning balances task accuracy against direct disclosure and registry-based re-identification risk, with privacy evaluated over the complete outbound transcript of each encounter. A local lesson memory distills completed consultations into generalized clinical guidance and retrieves relevant lessons before transmission, allowing subsequent cases to reuse expertise without another remote exchange. Memory grows without additional outcome labels or parameter updates. On an emergency-department benchmark built from MIMIC-IV-ED records, PrivMeSA improves mean task accuracy over delegation by up to 15.8 percentage points. In the same setting, PrivMeSA reduces the disclosure of personal details from 98.0% to 0.2% of cases and the share of cases in which the patient can be narrowed to ten or fewer registry patients from 74% to 0%.
Figures & tables
Figure 1: Re-identification in multi-turn clinical delegation. QIs that an attacker extracts from the consultations of the Delegation baseline narrow a registry of 205,504 patients to 9.
Figure 2: Overview of PrivMeSA. The local doctor drafts consultations, which the memory gate answers with a stored lesson or a miss. The doctor then uses the lesson, continues locally, or revises the draft and sends it to a remote specialist. An attacker counts the registry patients matching the transcript’s quasi-identifiers, which sets the anonymity term of the GRPO reward, and completed consultations are distilled into lessons for later cases.
Accuracy ↑
Disclosure
k≤10↓
Method
Disposition
Diagnosis
Procedure
Mean
Leak ↓
Per visit
Linked
Local model: Gemma 4 12B
Delegation
0.701
0.340
0.650
0.564
94.3%
48%
92%
PAPILLON
0.616
0.340
0.623
0.526
21.2%
1%
8%
PrivMeSA
0.692
0.380
0.705
0.592
0.9%
0%
0%
Local model: Granite 4.2 8B
Table 1: Main results. Mean averages the three task accuracies. Leak is the share of cases that sent any injected personal detail. k counts registry patients matching the facts sent, per visit and after linking a patient’s visits. Best value per local model in bold.
Direct disclosure ↓
Re-identification
Per visit
Linked
Method
Job
Employer
Kin
Address
Any
Ident. ↓
Med. k↑
k≤10↓
Med. k↑
k≤10↓
Local model: Gemma 4 12B
Delegation
69.3%
61.8%
72.8%
42.4%
94.3%
5.96
12
48%
1
92%
PAPILLON
21.2%
0.0%
0.0%
0.0%
21.2%
2.14
7,265
1%
1,183
8%
PrivMeSA
0.9%
0.0%
0.0%
0.0%
0.9%
1.35
7,916
0%
3,588
0%
Table 2: Privacy results. Direct disclosure: share of cases in which each injected personal detail reached a remote model; Any also counts the patient’s full name. Re-identification: identifiers an attacker recovers per visit, and k , the number of registry patients matching them, for each visit alone and after linking all of a patient’s visits. Best value per local model in bold.
Accuracy ↑
Per-visit median k↑
Linked median k↑
Model
Q1
Q2
Q3
Q4
Q1
Q2
Q3
Q4
Q1
Q2
Q3
Q4
PAPILLON
0.526
0.525
0.517
0.541
7,226
7,343
7,520
7,476
3,472
794
1,902
1,347
PrivMeSA, memory off
0.557
0.620
0.600
0.579
7,277
7,780
7,371
7,459
3,182
3,558
3,625
3,176
PrivMeSA, memory on
0.558
0.628
0.606
0.581
7,746
7,662
7,774
7,517
3,363
3,596
3,548
3,584
Table 3: Self-evolving memory by quarter of the test sequence using Gemma 4 12B as local model. “Memory off” keeps the memory empty. “Memory on” starts empty and grows after every case. Accuracy is the mean over three tasks. Median k is the median number of registry patients matching the facts sent, for each visit alone and after linking all of a patient’s visits.
Accuracy ↑
Disclosure
k≤10↓
Method
Disposition
Diagnosis
Procedure
Mean
Leak ↓
Per visit
Linked
PrivMeSA
0.692
0.380
0.705
0.592
0.9%
0%
0%
PrivMeSA (Base Gemma)
0.648
0.369
0.645
0.554
90.3%
33%
90%
PrivMeSA (no probing)
0.704
0.364
0.695
0.588
80.8%
17%
81%
Local only
0.655
0.208
0.591
0.485
–
–
–
Local only + RL
0.713
0.305
0.668
0.562
–
–
–
Table 4: Ablation studies. Each variant is evaluated with the same specialists. Base Gemma is the untrained policy, and no probing trains against specialists that never request personal information. Local-only variants never consult. Best value in bold.
Figure 3: Partial exchanges of consultations of the three systems on the diagnosis and procedure tasks of one test visit, an 81-year-old woman with dyspnea who was admitted with no coded procedure. Diagnoses are graded by CCS category.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Action
Effect
get_vitals
Return vital signs recorded since the previous call
get_medications
Return home and administered medications recorded so far
retrieve_past_visits
Return summaries of the most recent earlier visits
search_orders
Search the order catalog
order
Order laboratory tests, imaging, microbiology or medications
wait
Advance to the next recorded event
Appendix
Table 5: Actions available to the local agent.
Task
Label source
Eligible visits
Disposition (home or not)
ED disposition
Home, admitted or transferred
Diagnosis (CCS category)
First-listed ED diagnosis
Diagnosis with a CCS category
Procedure (any coded)
Hospital procedure codes
Linked hospital admission
Appendix
Table 6: Task definitions on the main test set.
Per visit
Linked
Attacker
Identifiers
Median k
k≤10
Median k
k≤10
GPT-5.6-Luna (evaluation)
1.35
7,916
0%
3,588
0%
Gemma 4 12B (training)
1.32
7,578
0%
3,758
0%
Appendix
Table 7: Re-identification of PrivMeSA under the evaluation and training attackers. Identifiers is the mean number of quasi-identifiers extracted per visit.
Task
Diagnosis
Key
ICD code for an adult presenting with influenza without specific occupational or environmental exposures.
Lesson
Use J11.1 (Influenza due to unidentified influenza virus with other respiratory manifestations) as the primary ED diagnosis for uncomplicated influenza when no specific strain or work-related exposure is documented.
Task
Procedure
Key
Whether an operative procedure will be coded for a patient with constipation and normal electrolytes.
Lesson
No operative procedure is typically coded for uncomplicated constipation with normal electrolytes and no evidence of obstruction or perforation. Routine medical management is expected unless a specific procedure like an enema or manual disimpaction is explicitly documented.
Appendix
Table 8: An example of two actual lessons from the PrivMeSA lesson memory.
Setting
Gemma 4 12B
Granite-4.2-8B
LoRA α / dropout
64 / 0
64 / 0
Rollout temperature
1
1
LoRA target modules
attention and MLP projections
attention and MLP projections
Optimizer
AdamW, weight decay 0.01
AdamW, weight decay 0.01
Learning-rate schedule
10-step warmup, linear decay
10-step warmup, linear decay
Training steps
200
120
Appendix
Table 9: Training settings for the two local models.
Figure 4: Test-time learning with Gemma 4 12B. Memory on starts empty and grows after every case, whereas memory off keeps it empty. (a) Lessons in memory. (b) Share of consultation drafts answered from memory. (c) Messages sent to the specialists per case. (d) Quasi-identifiers extracted by the attacker per case. Panels (b) to (d) are rolling means over 225 cases.
Per visit
Linked
Reward
Accuracy ↑
Leak ↓
Ident. ↓
Med. k↑
k≤10↓
Med. k↑
k≤10↓
Anonymity per visit or linked (test set)
Per-visit anonymity
0.583
0.0%
3.36
730
13%
3
61%
Linked anonymity
0.584
0.0%
6.34
9
55%
1
96%
Privacy terms or accuracy alone (training episodes)
Full reward
0.502
6.6%
3.55
1,954
–
–
–
Appendix
Table 10: Reward ablations with Gemma 4 12B. Top: anonymity rewarded per visit or over a patient’s first three visits linked, evaluated on the test set with the evaluation attacker. Bottom: the full reward or accuracy alone, measured on training episodes by the training attacker and averaged over steps 129 to 148, with median k per visit. Best value per panel in bold.
Task
Prompt
Answer
Disposition
Predict the recorded final ED disposition for this visit.
home, admitted, transfer
Diagnosis
Give the single ICD code for the primary ED diagnosis that will be recorded for this visit. Answer with the code only, in either ICD-9-CM or ICD-10-CM.
free text
Procedure
This patient is admitted. Predict whether any procedure will be coded for the hospital stay.