Strong Multilingual Privacy Tagging at Encoder Speed
Organizations: RWS Language Weaver
Abstract
Privacy redaction must remove personal information while preserving relationships expressed in text. We develop a multilingual named-entity tagger with fine-grained distinctions supporting varied redaction policies and methods for cheaply learning additional distinctions. We fine-tune a multilingual encoder with an affine span-tagging head on frontier-model annotations in 35 languages, replay mapped human gold with coverage-aware masking so unannotated types are not treated as negatives, and repair subword boundaries with a learned +/-1-character adjustment. On 1,283 human-gold test segments in seven languages, best measured redaction F1 is 88.8, against 69.1 for published GLiNER2 with 11 unrepresentable types excluded from its task (68.8 without that exemption), 67.8 for GLiNER2 adapted to the new training data, 57.3 for Microsoft Presidio and 35.8 for the best published OpenAI Privacy Filter fine-tune. Adding about 50,000 annotated training sentences and increasing human-gold replay improves exact typed-span F1 from 74.5 to 76.3 on Ont3, our 31-type frontier-annotated NER evaluation of 1,201 development segments. Mapped-gold replay alone raises human-gold F1 by ten points without loss on frontier-annotated text; boundary adjustment adds 1.7 exact typed-span F1 points on Ont3. Local LLMs fitting on a single 96-GB GPU underperformed as prompted annotators and frozen encoders, with encoding 30-95 times slower than XLM-R inference and prompted annotation roughly 180-1,100 times slower in the evaluated configurations. The encoder architecture delivers 4.9 times GLiNER2's CPU throughput. We release code, prompts and training recipes, with data-acquisition scripts and source links.
Figures & tables
| Recipe | Unique non-human texts | Labeled tokens | Human-gold draws | Input |
|---|---|---|---|---|
| O3 | 38,700 | 477,685 | 40% | Neighboring sentences |
| O4 | 88,512 | 856,116 | 50% | Isolated sentence |
| System | Hardware | Input tokens/s |
|---|---|---|
| XLM - R | CPU | 675 |
| GLiNER2 | CPU | 138 |
| XLM - R | L40S | 2,744 |
| GLiNER2 | L40S | 1,405 |
| Qwen3.8-Flash - Next | RTX PRO 6000 | 29 |
| Gemma - 4-31B base | RTX PRO 6000 | 93 |
Appendix figures & tables44 assets
Supplementary material from the paper’s appendix.
Appendix
| Human-gold component | Inputs | Source labels used by the mapping | Accepted primary types | Gold spans with multiple accepted types / annotated spans |
|---|---|---|---|---|
| MAPA | 189 | 6 | 13 | 59 / 126 |
| AQMAR | 134 | 3 | 8 | 82 / 173 |
| OpenNER collection | 904 | 3 | 8 | 350 / 999 |
| Wojood subset | 56 | 11 | 14 | 52 / 99 |
| Resource | Terms and access |
|---|---|
| Nemotron-PII , OpenPII 1M and 1.5M | CC BY 4.0; public downloads. |
| Ai4Privacy 200k | Custom terms; for-profit organizations with more than three staff need a corporate license. Redistribution restrictions apply. |
| Ai4Privacy health-PHI 400k diagnostic sample | Gated access under custom commercial terms. |
| TAB , idner-news-2k and Wojood public sample | MIT; public repositories. Full Wojood is separate and was not used. |
| MEDDOCAN , SPY , MultiGraSCCo and MAPA EUR-Lex | CC BY 4.0; public distributions. Other MAPA packages may differ. |
| KLUE NER and HiNER | CC BY-SA 4.0; public distributions. |
| Training scope | Updates | GPU / AWS instance | Hours | Estimated $ |
|---|---|---|---|---|
| Ont2-equivalent exposure | 33,300 | L40S / g6e.4xlarge | 6.0* | 8.5–18.0* |
| Ont3 fit from pretrained XLM - R | 16,000 | RTX PRO 6000 Blackwell Server / g7e.2xlarge | 2.3 | 4.3–5.4 |
| Model | Failed format (%) | Character F1 | With XLM - R fallback* | Rescued subset F1* |
|---|---|---|---|---|
| Gemma - 4-12B | 5.01 | 83.84 | 86.38* | 84.51* |
| Gemma - 4-31B-FP8 | 6.22 | 83.98 | 90.11* | 85.13* |
| Qwen3.8-Flash - Next-NVFP4 | 8.19 | 76.03 | 87.02* | 90.77* |
| O4, sentence only | — | 91.53 | — | — |
| System | Params (B) | Active (B) | Labeled input tokens/s |
|---|---|---|---|
| Gemma - 4-12B | 12 | — | 2 |
| Gemma - 4-31B-IT, bfloat16 | 30.7 | — | 2.5 |
| Qwen3.8-Flash - Next-NVFP4 | 176 | 6 | 15 |
| XLM - R encoder | 0.55 | — | 2,744 |
| (Prompted) LLM / encoder | Params (B) | Head | Input state offsets | Typed span F1 (%) | Redaction-character F1 (%) |
|---|---|---|---|---|---|
| Flash-Next unprompted | 176 | Affine | 0 | 39.14 | 78.59 |
| Qwen3.8-Flash-Next | 176 | Affine | 0 | 44.56 | 80.64 |
| Qwen3.8-Flash-Next | 176 | Affine | −1, 0, +1 | 53.00 | 82.96 |
| Qwen3.8-Flash-Next | 176 | 512-unit GELU | 0 | 53.80 | 85.25 |
| Qwen3.8-Flash-Next | 176 | 512-unit GELU | −1, 0, +1 | 59.54 | 86.61 |
| Gemma - 4-31B base | 30.7 | 512-unit GELU | −1, 0, +1 | 62.41 | 87.07 |
| Coverage | Importance | Language codes |
|---|---|---|
| Original 20 | ×4 | en |
| ×2 | ar, de, es, fr, ko, pt, vi, zh | |
| ×1 | cs, hi, id, it, ja, nl, pl, ru, sv, tr, uk | |
| +15 → 35 | ×1 | bn, da, el, fa, fi, fil, he, hr, ms, no, ro, ta, te, th, ur |
| Updates | Human gold exact regions | Ont3 exact regions |
|---|---|---|
| 0 (transferred initialization) | 57.44 | 50.11 |
| 100 | 63.73 | 54.84 |
| 200 | 64.72 | 56.96 |
| 300 | 64.52 | 56.24 |
| 400 | 62.98 | 57.10 |
| 600 (selected) | 67.24 | 54.79 |
| Baseline configuration | Languages | Documents | Our model’s F1 gain over baseline (points) | 95% interval (points) |
|---|---|---|---|---|
| GLiNER2 | 7 | 421 | +21.78 | [+19.14, +24.42] |
| OpenAI Privacy Filter | 1 | 61 | +47.31 | [+42.37, +52.25] |
| Multilingual OpenMed | 14 | 907 | +40.55 | [+37.63, +43.46] |
| OpenMed with additional task fine-tuning | 20 | 1,267 | +48.70 | [+46.08, +51.31] |
| OpenMed Nemotron | 9 | 541 | +43.76 | [+40.18, +47.35] |
| Method | O2 | O3 | O4 |
|---|---|---|---|
| Starting encoder | Previously task-trained encoder, transferred to Ont2 | Upstream pretrained XLM - R-large | Ont3 encoder’s isolated-input branch |
| Language coverage | 20 languages | 35 languages | 35 languages |
| Supervision | Mapped annotations, translated data and Ont2 complete-label calibration mixture | Ont3 mixture plus 40% draws from four mapped human-gold corpora | 49,812 additional frontier-annotated texts; earlier labels cleaned; prompts clarify occupation/job titles as demographic_attribute , outside names. 50% mapped human-gold draws. |
| Primary label inventory | 29 types | Same 29 types plus two optional reference types | Same as O3 |
| Additional training objectives | Binary entity-presence loss, weight 2 | Subclass and type-conditioned predicate losses, each weight 1 | Same as O3 |
| O loss weight | 1 on complete annotations; 0 on partial annotations | 0.75 on supervised O tokens; corpus-unannotated types masked | Same as O3 |
| Decoder | Typed exact F1 (%) | Typed overlap F1 (%) | CPU decode time (s) |
|---|---|---|---|
| Matching-type Viterbi | 73.77 | 78.65 | 2.41 |
| Coarse-compatible greedy; first label | 70.51 | 75.61 | 0.32 |
| Coarse-class Viterbi; max scores | 74.00 | 78.91 | 1.54 |
| Coarse-class Viterbi; soft scores, T = 0.5 | 74.19 | 79.18 | 2.33 |
| Score-adjusted Viterbi; α = 0.95, T = 0.5 | 74.19 | 79.17 | 2.53 |
| Token features and head | Head params | Initial screen F1 (%) | Repeat screen F1 (%) | Interpretation |
|---|---|---|---|---|
| final layer [24], affine | 357,725 | 82.20 | 82.14 | stable reference |
| middle + final [12,24], affine | 715,101 | 83.82 | 81.86 | initial gain does not repeat |
| layers [23,24], affine | 715,101 | — | 73.73 step 1,250 | early-screen loss |
| layers [8,24], affine | 715,101 | — | 71.67 step 1,250 | early-screen loss |
| embedding + final [0,24], affine | 715,101 | — | 69.16 step 1,250 | early-screen loss |
| final layer, 1,024-unit GELU | 1,409,373 | 69.21 step 1,250 | — | early-screen loss |
| Encoder and task head | Fresh6 F1 (%) [95% interval] | MultiGraSCCo F1 (%) [95% interval] | Disposition |
|---|---|---|---|
| XLM - R-large, final affine | 92.23 [91.41, 92.98] | 75.07 [74.27, 75.90] | quality base selected |
| LaBSE, final affine | 87.88 [86.68, 88.94] | 72.53 [71.69, 73.38] | common-head comparison |
| mmBERT-base, stock head | 89.25 [88.27, 90.17] | 70.61 [69.77, 71.49] | faster challenger; different head |
| Model and head | Human gold regions | Ont3 regions | Ont3 typed |
|---|---|---|---|
| mmBERT-base, stock GELU | 84.89 | 76.86 | 72.32 |
| mmBERT-base, affine | 87.03 | 77.02 | 73.90 |
| XLM - R-large O4, affine | 87.80 | 80.20 | 76.30 |
| Seed | Name repair: macro span Δ (points) [95% interval] | Name repair: macro fine Δ (points) [95% interval] |
|---|---|---|
| 154 | +1.57 [+0.79, +2.38] | +2.47 [+1.78, +3.20] |
| 155 | +1.44 [+0.64, +2.23] | +1.96 [+1.30, +2.62] |
| 156 | +1.88 [+0.92, +2.84] | +2.64 [+1.86, +3.42] |
| Seed | Date repair: Korean span Δ (points) [95% interval] | Date repair: macro span Δ (points) [95% interval] |
|---|---|---|
| 154 | +5.08 [+2.64, +7.63] | +1.62 [+0.92, +2.33] |
| 155 | +7.53 [+4.61, +10.63] | +2.12 [+1.29, +2.96] |
| 156 | +4.38 [+1.06, +7.58] | +1.16 [+0.27, +2.04] |
| Step | Text |
|---|---|
| Source field | First Name: Kevin |
| Replace the name with a placeholder | First Name: [GIVEN_NAME_1] |
| TranslateGemma output | Nombre: [GIVEN_NAME_1] |
| Surface | View | XLM - R Δ (F1 points) [95% interval] | mmBERT Δ (F1 points) [95% interval] |
|---|---|---|---|
| Fresh20 | span overlap | +0.59 [+0.13, +1.06] | +0.44 [+0.18, +0.71] |
| Fresh20 | fine overlap | +0.20 [-0.08, +0.49] | +0.30 [+0.08, +0.53] |
| MultiGraSCCo | span overlap | +0.15 [+0.04, +0.26] | +0.38 [+0.22, +0.53] |
| MultiGraSCCo | fine overlap | +0.05 [-0.04, +0.13] | +0.44 [+0.28, +0.61] |
| Evaluation | Untyped spans | Fine labels |
|---|---|---|
| Fresh20 (twenty-language development set) | +1.36 [0.90, 1.82] | −0.45 [−1.02, 0.10] |
| MultiGraSCCo | +0.46 [0.24, 0.68] | −0.90 [−1.19, −0.61] |
| Stage | Accepted at this stage | Cumulative acceptance |
|---|---|---|
| TranslateGemma-27B, after marker repair | 68,309 / 69,480 (98.31%) | 98.31% |
| Gemma - 4-31B retry of the 1,171 failures | 1,007 / 1,171 (86.00%) | 99.76% |
| Model | Format passed | Format and replacement type passed |
|---|---|---|
| TranslateGemma-12B | 0% (0/5) | 0% (0/5) |
| Prompted Gemma - 4-12B | 100% (5/5) | 60% (3/5) |
| Model | Before repair | After numeric-ID recovery |
|---|---|---|
| TranslateGemma-12B | 25.0% (31/124) | 40.3% (50/124) |
| Weighted view | Mean delta (points) | Paired 95% interval (points) | Adjusted status |
|---|---|---|---|
| span overlap | +1.00 | [+0.73, +1.40] | win |
| character precision | -0.26 | [-0.38, -0.17] | loss |
| character recall | +2.80 | [+2.40, +3.20] | win |
| character F1 | +1.50 | [+1.20, +1.70] | win |
| fine overlap | +0.11 | [-0.20, +0.42] | unresolved |
| Annotator | Span density relative to Luna | Without _reference tags |
|---|---|---|
| Gemma - 4-31B | 82% | 90% |
| Qwen3.8-27B | 61% | 75% |
| Model | Languages | Sentences | Exact typed span F1 (%) | Character-redaction F1 (%) |
|---|---|---|---|---|
| Ont2 | Shared 20 | 374 | 52.52 | 76.63 |
| Ont3 | Shared 20 | 374 | 73.21 | 91.58 |
| Ont2 | 35 | 659 | 51.33 | 75.09 |
| Ont3 | 35 | 659 | 72.89 | 91.17 |
| Population | Exact-match task | Without model | With model | ΔF1 (95% interval) |
|---|---|---|---|---|
| Human gold, 1,283 segments | Redaction regions | 87.87 | 87.80 | −0.07 [−0.23, 0.00] |
| Ont3, 1,201 segments | Redaction regions | 77.71 | 80.20 | +2.49 [+1.58, +3.50] |
| Ont3, 1,201 segments | Fine typed spans | 74.62 | 76.30 | +1.68 [+1.01, +2.41] |
| Prompt | F1 (%) | ΔF1 (points) | Paired 95% CI | Time (s) |
|---|---|---|---|---|
| No prompt | 72.53 | 0 | — | 2.64 |
| 10 shared tokens | 71.65 | −0.880 | [−2.133, +0.394] | 2.59 |
| 8 shared + 2 language tokens | 71.94 | −0.587 | [−2.012, +0.848] | 2.59 |
| 58 learned tokens, 29 at inference | 73.16 | +0.63 | — | 2.91 |
| Gold status: 58 active in train, 29 at inference | 73.34 | +0.807 | [−0.514, +2.210] | 2.95 |
| O4: no prompt | 75.74 | 0 | — | — |