Medical images are released with the reports that describe them, and protecting the image does not protect the report. This paper measures the text component of such releases. We measure identifier detection, downstream utility and residual identity leakage on the same documents, with the pseudonymisation policy as the variable under test: 15 detectors, three release conditions and four corpora of medical reports, legal judgments, news and other genres, and e-mail, in German, English, Chinese and Arabic. A fixed 13-detector union reaches a person sensitivity of 0.9998 at specificity 0.8686 on the medical reports, 0.9958 at 0.8504 on the legal judgments, 0.9352 at 0.9318 on news and other genres, and 0.9906 at 0.6235 on e-mail. With this ensemble, frequency matching with a public name list recovers zero identities by alignment across the four corpora; the names it got right were ones the detector missed, left in clear text. Cross-document linkage ranks the correct person first for 0.93% of e-mail queries without training and 3.94% with it, against 1/3697 chance and 71.98% on unmodified text. On the medical reports it recovers nothing without training and 0.71% of 138 queries with it, against 1/207 chance and a 2.73% ceiling on unmodified text.
Figures & tables
Corpus
Docs
Tokens
Ment./doc
Lang.
Reference annotation
CARDIO:DE
400
885059
78.3
de
constructed
TAB
1268
1844804
81.2
en
manual
OntoNotes
5994
4726594
45.5
en/zh/ar
linguistic
Enron
58636
15507925
12.2
en
header-derived
Table 1 : The corpora separate domain, language, and where the reference annotation came from.
CARDIO:DE
TAB
OntoNotes
Enron
Operating point
Sens.
Spec.
Cost
Sens.
Spec.
Cost
Sens.
Spec.
Cost
Sens.
Spec.
Cost
max. sensitivity
0.9919
0.9670
57.633∪
0.9253
0.8986
14.210∪
0.7466
0.9763
9.450∪
0.9968
0.7745
25.557∪
fast sensitivity
0.8947
0.9769
0.301∪
0.8926
0.9242
0.219∪
0.6843
0.9731
5.376∪
0.9222
0.7420
0.057∪
max. specificity
0.5230
0.9997
116.875v
0.5014
0.9959
98.110v
0.5034
0.9960
0.103∪
0.5292
0.9448
20.931∩
fast specificity
0.7807
0.9862
0.263s
0.5198
0.9934
0.117s
0.5034
0.9960
0.103∪
0.6920
0.8781
0.026s
Table 2 : Operating points on both error rates. Sensitivity is over the identifier types the conditions replace, specificity over every other token. Cost is seconds per document; the rule is ∪ union, ∩ intersection, v at-least- k vote, s a single detector.
Corpus
Sens.
Spec.
Exp. (%)
Frequency
Context link.
Learned link.
Chance
CARDIO:DE
0.9998
0.8686
0.05
0 (1)
0
0.71±1.60
1/207
Enron
0.9906
0.6235
1.60
0 (0)
0.93±0.10
3.94±0.30
1/3697
TAB
0.9958
0.8504
0.83
0 (1)
no links → not measurable
–
OntoNotes
0.9352
0.9318
5.43
0 (1)
no links → not measurable
–
Table 3 : Leakage from the recommended 13-detector union; all columns concern people. Sens. is person sensitivity and Spec. the share of non-identifier tokens left alone: neither is readable without the other. Exp. is the risk that a person is still named in a given document after processing, Frequency the identities named from frequency alone with a public list and, bracketed, with the corpus’s own distribution (an upper bound, not a risk). Linkage is the percentage of queries ranking the correct person first, over the complete candidate list, against chance 1/gallery .
Adrem Data Lab, Department of Computer Science, University of Antwerp, Antwerp, Belgium · Antwerp University Hospital (UZA), Edegem, Belgium · Laboratory of Experimental Medicine and Pediatrics (LEMP), University of Antwerp, Antwerp, Belgium
University of Notre Dame, Notre Dame, IN, USA. · IBM Research, Yorktown Heights, NY, USA. · Department of Cardiology, Boston Children’s Hospital, Boston, MA, USA. +2